TranscriptBalance turns audio and video files into text for free and without limits, right in your browser. The Whisper speech recognition model runs on your own device, so your files are never uploaded and you don’t need an account.
Audio and video files, for example MP3, M4A, WAV, MP4 or MOV, also several at once. You can save the text as TXT, Word, Markdown, JSON or CSV, and subtitles as SRT or VTT. Several files come as one document or as separate files in a ZIP.
Whisper knows 99 languages, including English, German, French, Dutch, Spanish, Italian, Polish and Turkish. TranscriptBalance detects the language automatically, or you pick it yourself. Widely spoken languages work best.
It depends on the model, the recording and the language. You can choose from four Whisper models, from “Flash” to “Precise” (large-v3-turbo, GPU only), and the app suggests the one that suits your device. With clear speech the larger models usually produce very good text; accents, jargon, background noise and people talking over each other cause more mistakes. Always check names, numbers and quotes.
Yes, if you turn on “Speakers”. Speaker detection also runs on your device, automatically or for 2 to 6 people, and you can rename the speakers. When two people talk at once, a sentence may be assigned to the wrong person.
Yes. The AI summary of a lecture, meeting, interview or anything else is 1 to 5 pages long. It is also made on your device: with Chrome’s built-in AI or with Gemma (needs WebGPU).
There is no fixed limit. The audio is processed in your device’s memory, so a computer handles longer recordings than a phone. If memory runs out, split the recording into shorter files. Keep the tab open until everything is done.
The text appears while you listen: from a browser tab, from your screen’s audio (such as the Teams app) or from the microphone. Tab and screen audio work in Chrome and Edge on a computer; the microphone also works on a phone. Live mode uses smaller models so it can keep up. The audio is never stored.
Ask the AI chat (Chrome AI, or Gemini with your own key) things like “What did I miss?”, and it answers questions that come up in the conversation by itself. If you switch it on, a summary grows every 5 minutes. When you stop, the AI writes an overview with appointments and to-dos, and you can add the dates it found to your calendar in one click.
The transcript itself is always made on your device. Chrome AI also runs on your device; it is only available in Google Chrome on a computer, and only if your device supports it. If you choose to connect Gemini with your own key, your browser sends excerpts of the transcript and your questions straight to Google, not to us. You need a Google account for that, Google only allows it from age 18, and Google’s terms and limits apply.
Ask the lecturer first, tell the others in a meeting, get their consent where it is required, and follow the rules of your university or employer. TranscriptBalance never stores audio, but the transcript contains what other people say. Laws differ by country: in Germany, for example, recording someone’s private spoken words without consent can be a criminal offence. Don’t publish lecture transcripts without permission.
Only text is saved, never the audio: every session is kept in your browser under “Saved sessions”, sorted by subject if you like, and you can keep asking questions about it later. You can download the transcript, the overview and the live summary as Word, TXT or Markdown.
No. Your audio never leaves your device: your browser reads files locally, speech recognition runs on your device, and transcripts are kept only in your browser’s storage. The only things fetched from the internet are the models (from Hugging Face) and the runtime that runs them (from jsDelivr). Only if you choose to connect Gemini in live mode do excerpts of the transcript and your questions go straight to Google.
Yes. There is no subscription, no paid tier, no third-party advertising and no limit on minutes or files; the only limit is your device, as long recordings take time and memory. Because your device does the work, TranscriptBalance needs no expensive servers. It is built by Jeremy Heindrichs, a student in Belgium, alongside his studies. If you’d like to support him, take a look at his other project, AudioBalance: a Windows app that keeps videos, calls and music at the same volume.
No. TranscriptBalance works without an account, an email address or cookies. Only if you choose to use Gemini in live mode do you need your own Google account, your own API key from Google and to be 18 or older.
We keep an anonymous count of visits and features used, never with content and without cookies. You can switch off the messages about features used in the app under “About & licenses”. With Global Privacy Control or Do Not Track your visit is not counted either. The details are in the privacy policy.
It is fastest with WebGPU, for example in a current version of Chrome or Edge on a computer. Without WebGPU, TranscriptBalance runs on the processor: slower, but in practically any current browser, including on phones and tablets. Capturing audio from a tab or your screen only works in Chrome or Edge on a computer.
Not completely. You need an internet connection to open the page and for the first model download. After that the models are stored in your browser and the transcription runs on your device, but there is no dedicated offline mode.