Transcribe audio and video to text for free – no upload, no sign-up

TranscriptBalance turns audio and video files into text for free and without limits, right in your browser. The Whisper speech recognition model runs on your own device, so your files are never uploaded and you don’t need an account.

Transcribing files

How do I turn a file into text?
  1. Add: drag one or more files into TranscriptBalance, or click the round area in file mode.
  2. Transcribe: the first time, your browser downloads the speech model and keeps it. The work happens on your device, on the graphics card (WebGPU) or the processor.
  3. Download: save the text in the format you need, with or without timestamps and speaker names.
Which formats does TranscriptBalance support?

Audio and video files, for example MP3, M4A, WAV, MP4 or MOV, also several at once. You can save the text as TXT, Word, Markdown, JSON or CSV, and subtitles as SRT or VTT. Several files come as one document or as separate files in a ZIP.

Which languages does TranscriptBalance support?

Whisper knows 99 languages, including English, German, French, Dutch, Spanish, Italian, Polish and Turkish. TranscriptBalance detects the language automatically, or you pick it yourself. Widely spoken languages work best.

How accurate is the transcription?

It depends on the model, the recording and the language. You can choose from four Whisper models, from “Flash” to “Precise” (large-v3-turbo, GPU only), and the app suggests the one that suits your device. With clear speech the larger models usually produce very good text; accents, jargon, background noise and people talking over each other cause more mistakes. Always check names, numbers and quotes.

Can TranscriptBalance tell speakers apart?

Yes, if you turn on “Speakers”. Speaker detection also runs on your device, automatically or for 2 to 6 people, and you can rename the speakers. When two people talk at once, a sentence may be assigned to the wrong person.

Can TranscriptBalance summarize the text?

Yes. The AI summary of a lecture, meeting, interview or anything else is 1 to 5 pages long. It is also made on your device: with Chrome’s built-in AI or with Gemma (needs WebGPU).

How large or long can a file be?

There is no fixed limit. The audio is processed in your device’s memory, so a computer handles longer recordings than a phone. If memory runs out, split the recording into shorter files. Keep the tab open until everything is done.

Live transcription

How does live mode work?

The text appears while you listen: from a browser tab, from your screen’s audio (such as the Teams app) or from the microphone. Tab and screen audio work in Chrome and Edge on a computer; the microphone also works on a phone. Live mode uses smaller models so it can keep up. The audio is never stored.

What does the AI do in live mode?

Ask the AI chat (Chrome AI, or Gemini with your own key) things like “What did I miss?”, and it answers questions that come up in the conversation by itself. If you switch it on, a summary grows every 5 minutes. When you stop, the AI writes an overview with appointments and to-dos, and you can add the dates it found to your calendar in one click.

Chrome AI or Gemini: where does my data go?

The transcript itself is always made on your device. Chrome AI also runs on your device; it is only available in Google Chrome on a computer, and only if your device supports it. If you choose to connect Gemini with your own key, your browser sends excerpts of the transcript and your questions straight to Google, not to us. You need a Google account for that, Google only allows it from age 18, and Google’s terms and limits apply.

Am I allowed to transcribe lectures or Teams meetings live?

Ask the lecturer first, tell the others in a meeting, get their consent where it is required, and follow the rules of your university or employer. TranscriptBalance never stores audio, but the transcript contains what other people say. Laws differ by country: in Germany, for example, recording someone’s private spoken words without consent can be a criminal offence. Don’t publish lecture transcripts without permission.

Are live sessions saved, and can I export them?

Only text is saved, never the audio: every session is kept in your browser under “Saved sessions”, sorted by subject if you like, and you can keep asking questions about it later. You can download the transcript, the overview and the live summary as Word, TXT or Markdown.

Privacy, cost and requirements

Are my recordings uploaded?

No. Your audio never leaves your device: your browser reads files locally, speech recognition runs on your device, and transcripts are kept only in your browser’s storage. The only things fetched from the internet are the models (from Hugging Face) and the runtime that runs them (from jsDelivr). Only if you choose to connect Gemini in live mode do excerpts of the transcript and your questions go straight to Google.

Is TranscriptBalance really free and unlimited?

Yes. There is no subscription, no paid tier, no third-party advertising and no limit on minutes or files; the only limit is your device, as long recordings take time and memory. Because your device does the work, TranscriptBalance needs no expensive servers. It is built by Jeremy Heindrichs, a student in Belgium, alongside his studies. If you’d like to support him, take a look at his other project, AudioBalance: a Windows app that keeps videos, calls and music at the same volume.

Do I need to sign up?

No. TranscriptBalance works without an account, an email address or cookies. Only if you choose to use Gemini in live mode do you need your own Google account, your own API key from Google and to be 18 or older.

What does TranscriptBalance count, and can I turn it off?

We keep an anonymous count of visits and features used, never with content and without cookies. You can switch off the messages about features used in the app under “About & licenses”. With Global Privacy Control or Do Not Track your visit is not counted either. The details are in the privacy policy.

Which browsers and devices does it work on?

It is fastest with WebGPU, for example in a current version of Chrome or Edge on a computer. Without WebGPU, TranscriptBalance runs on the processor: slower, but in practically any current browser, including on phones and tablets. Capturing audio from a tab or your screen only works in Chrome or Edge on a computer.

Does TranscriptBalance work offline?

Not completely. You need an internet connection to open the page and for the first model download. After that the models are stored in your browser and the transcription runs on your device, but there is no dedicated offline mode.