TranscriptBalance
Open app

Transcribe research interviews for free – for your thesis or dissertation

Updated

TranscriptBalance transcribes interviews for your thesis or dissertation for free, with no minute limit, speaker detection, timestamps and export to Word. Your recording stays on your device: Whisper runs right in your browser, and nothing is uploaded.

Why transcribe interviews locally?

Research interviews often contain personal information: names, opinions, sometimes health or workplace details. Many transcription services upload your recording to their servers to do the work. TranscriptBalance transcribes in your browser, with no account. Neither the recording nor the transcript reaches a server, ours included, so no service processes your interviews on your behalf.

What does leave your device: the first time, your browser downloads the models from Hugging Face and the runtime from jsDelivr. These providers see your IP address, but no content. The app also counts anonymously which features are used, never content, and you can switch that off. Details are in the privacy policy and under transcription without upload.

Whether this is enough for your project depends on your university’s rules, your ethics approval, your data management plan and the consent form your participants signed. Check them or ask your supervisor or ethics committee. This is not legal advice.

From recording to transcript, step by step

  1. Record: with your phone, a voice recorder or by recording an online interview. Common audio and video files such as MP3, M4A, WAV or MP4 work.
  2. Drop the file: drag one or more recordings into TranscriptBalance.
  3. Set it up: choose a model, let the app detect the language or set it yourself, and switch on “Speakers”. The app works out the number of speakers, or you set 2 to 6.
  4. Transcribe: keep the tab open. With more than 10 minutes of audio, the app asks whether it may send you a notification when everything is done.
  5. Name the speakers: rename “Speaker A” and “Speaker B” to, say, “Interviewer” and “P1”, so no real names head the paragraphs.
  6. Export: as DOCX, TXT, SRT, VTT, Markdown, JSON or CSV, with or without timestamps and speakers. Several interviews go into one document or into a ZIP file, one file each.

Which model suits your interview?

The app recommends the model that suits your device. Each model is downloaded once, roughly 45 to 760 MB depending on the model and device, and then stays in your browser. For the transcript you’ll quote in your thesis, use the most accurate model your device can handle.

Tips for an accurate transcript

Long recordings and many interviews

There’s no limit on minutes or files. The app works through several recordings one after another and shows an estimated time remaining. How long it takes depends on your device and the model: with a graphics card via WebGPU it’s much faster than on the processor alone.

Keep the tab open; the app keeps your screen awake meanwhile. If the work is interrupted, for example because the tab was closed, the text so far stays saved in your browser. Drop the same file again and the app carries on after the last finished section. Very long files need a lot of memory, which can get tight on phones and low-end laptops.

If you like, an AI on your device writes an “Interview” summary of 1 to 5 pages: with Chrome AI, or with Gemma, which runs locally in the browser, needs WebGPU and is a one-time download of about 0.8 GB. In file mode nothing is sent to Gemini or any other online service. The summary is AI-generated; your analysis should rest on the transcript.

TranscriptBalance, aTrain or noScribe?

aTrain, from the University of Graz, and noScribe are well-known free programs that also transcribe locally with Whisper, detect speakers and are built for research interviews. The differences:

Which tool fits depends on your computer and the way you work. For a wider view, see the comparison of free transcription tools.

Transcribe an interview now

Frequently asked questions

Is TranscriptBalance GDPR-compliant for research interviews?

TranscriptBalance doesn’t upload your recording, needs no account and stores transcripts only in your browser; no service processes your interviews on your behalf. The only things that go online are the model downloads from Hugging Face and jsDelivr and anonymous usage counts without content, which you can switch off. Whether that meets your requirements is decided by your university’s or ethics committee’s rules. This is not legal advice.

How long does it take to transcribe a one-hour interview?

It depends heavily on your device and the model you choose. With a graphics card via WebGPU it’s much faster than on the processor alone, and smaller models are faster than large ones. TranscriptBalance shows an estimated time remaining while it works; it’s best to test with a short recording first.

Does TranscriptBalance detect who is speaking?

Yes. Speaker detection runs on your device, using the pyannote and WeSpeaker models. TranscriptBalance works out the number of speakers itself, or you set 2 to 6. You can then rename the speakers, for example to “Interviewer” and “P1”. Check the result where people overlap or sound alike.

Can I edit the transcript in Word?

Yes. TranscriptBalance exports a DOCX file, with or without timestamps and speaker names, that you correct in Word or any other word processor. TranscriptBalance has no built-in editor. TXT, SRT, VTT, Markdown, JSON and CSV are available too.

Which files can I transcribe?

Common audio and video formats such as MP3, M4A, WAV or MP4; for videos only the audio track is used. What works exactly depends on which formats your browser can read. If TranscriptBalance reports a format as unsupported, convert the file first, for example to MP3.

Is there a limit on length or number of interviews?

No. TranscriptBalance limits neither minutes nor files, and several recordings are processed one after another. The limits come from your device: very long files need a lot of memory, and slower devices take longer.

Are filler words and pauses transcribed?

Usually not. Whisper, the speech recognition behind TranscriptBalance, produces a cleaned-up text in which filler words like “um”, repetitions and pauses are often missing. If your transcription conventions call for a detailed transcript, add these by hand while proofreading.