Transcribe audio and video to text for free – no upload, no sign-up

TranscriptBalance turns audio and video files into text for free and without limits, right in your browser. The Whisper speech recognition model runs on your own device, so your files are never uploaded and you don’t need an account.

How it works

  1. Add your files: drag one or more into TranscriptBalance, or click the round area in file mode, for example MP3, M4A, WAV, MP4 or MOV.
  2. Transcribe: the first time, your browser downloads the speech model and keeps it. The work happens on your device, on the graphics card (WebGPU) or the processor.
  3. Download: save the text as TXT, Word, SRT, VTT, Markdown, JSON or CSV, with or without timestamps and speaker names.

What can TranscriptBalance do?

Live notes for lectures, seminars and Teams calls

In live mode the text appears while you listen: from a browser tab, from your screen’s audio (such as the Teams app) or from the microphone. Ask the AI chat (Chrome AI, or Gemini with your own key) things like “What did I miss?”, and it answers questions that come up in the conversation by itself. If you switch it on, a summary grows every 5 minutes, and when you stop, you can add the dates it found to your calendar in one click. The audio is never stored.

Tab and screen audio work in Chrome and Edge on a computer; the microphone also works on a phone. Read more: live lecture transcription.

Start live mode

Does my audio stay private?

Yes. Your audio never leaves your device, and transcripts are kept only in your browser’s storage. The only things fetched from the internet are the models (from Hugging Face) and the runtime that runs them (from jsDelivr). We also keep an anonymous count of visits and features used, never with content: you can switch off the feature messages in the app, and with Global Privacy Control or Do Not Track your visit is not counted either. Only if you choose to connect Gemini in live mode do excerpts of the transcript and your questions go straight to Google. How to check all this yourself: private transcription without uploads.

Why is it free?

Because your device does the work, TranscriptBalance needs no expensive servers. There is no subscription, no paid tier and no third-party advertising. It is built by Jeremy Heindrichs, a student in Belgium, alongside his studies. If you’d like to support him, take a look at his other project, AudioBalance: a Windows app that keeps videos, calls and music at the same volume.

See also: transcribing interviews for your thesis and free transcription tools compared.

Frequently asked questions

Is TranscriptBalance really free and unlimited?

Yes. There is no subscription, no minute limit and no cap on the number of files. The only limit is your device: long recordings take time and memory.

Is my audio file uploaded?

No. Your browser reads the file locally and Whisper transcribes it on your device. The only downloads are the models. All our server receives are anonymous counts without any content: you can switch off the messages about features used in the app under “About & licenses”, and with Global Privacy Control or Do Not Track your visit is not counted either.

Do I need to sign up?

No. TranscriptBalance works without an account, an email address or cookies. Only if you choose to use Gemini in live mode do you need your own Google account and your own API key from Google.

Which browsers and devices does it work on?

It is fastest with WebGPU, for example in a current version of Chrome or Edge on a computer. Without WebGPU, TranscriptBalance runs on the processor: slower, but in practically any current browser, including on phones and tablets. Capturing audio from a tab or your screen only works in Chrome or Edge on a computer.

How accurate is the transcription?

It depends on the model, the recording and the language. With clear speech the larger models usually produce very good text; accents, jargon, background noise and people talking over each other cause more mistakes. Live mode uses smaller models so it can keep up. Always check names, numbers and quotes.

Which languages does TranscriptBalance support?

Whisper knows 99 languages, including English, German, French, Dutch, Spanish, Italian, Polish and Turkish. TranscriptBalance detects the language automatically, or you pick it yourself. Widely spoken languages work best.

How large or long can a file be?

There is no fixed limit. The audio is processed in your device’s memory, so a computer handles longer recordings than a phone. If memory runs out, split the recording into shorter files. Keep the tab open until everything is done.

Can TranscriptBalance tell speakers apart?

Yes, if you turn on “Speakers”. Speaker detection also runs on your device, automatically or for 2 to 6 people, and you can rename the speakers. When two people talk at once, a sentence may be assigned to the wrong person.

Am I allowed to transcribe lectures or Teams meetings live?

That depends on the rules where you are. TranscriptBalance never stores audio, but the transcript contains what other people say. Tell them beforehand, get their consent where it is required, and follow the rules of your university or employer. Laws differ by country: in Germany, for example, recording someone’s private spoken words without consent can be a criminal offence. Don’t publish lecture transcripts without permission.

Does TranscriptBalance work offline?

Not completely. You need an internet connection to open the page and for the first model download. After that the models are stored in your browser and the transcription runs on your device, but there is no dedicated offline mode.