Private transcription without uploads: how Whisper runs in your browser
Updated
With TranscriptBalance, the Whisper speech recognition model runs right in your browser, on your own device. Your audio and video files are never uploaded: the only things that come from the internet are the models, downloaded once and then kept in your browser.
How does Whisper run in a browser?
Whisper is an open speech recognition model from OpenAI (MIT license). TranscriptBalance runs it with transformers.js and ONNX Runtime Web, two libraries that let AI models run inside the browser. There are two ways to do that:
- WebGPU: if your browser can use the graphics card, Whisper runs there. This is the fastest way.
- WebAssembly: without WebGPU, the same app runs on the processor, spread across several cores. This works in practically every current browser, but takes longer.
The app checks which way fits when it starts. If the graphics card cannot load a model or drops out, the app switches to the processor where possible. Your file is read in the browser, converted to 16 kHz audio, split at pauses into chunks of up to 28 seconds and transcribed chunk by chunk. Live mode always runs on the processor with small models, which keeps the graphics card free for the AI.
What is downloaded once?
Nothing when you open the page, only when you start a feature. Then your browser downloads:
- the Whisper model you picked, from Hugging Face. On the processor: Flash 44 MB, Fast 80 MB, Balanced 252 MB. With WebGPU: Flash 122 MB, Fast 209 MB, Balanced 589 MB, Precise (large-v3-turbo) 762 MB.
- speaker detection (pyannote and WeSpeaker, about 13 MB together), if you turn it on.
- for live mode, a dedicated German Whisper model (231 MB) or Whisper base for other languages (80 MB), plus Silero VAD (about 2 MB), which detects when someone is speaking.
- Gemma 3 1B (about 0.8 GB) for summaries, only if Chrome’s built-in AI cannot do the job.
- ONNX Runtime, the runtime, from jsDelivr, the first time a model starts.
Everything is stored in your browser (in the Origin Private File System) and loaded from there next time. The app also asks the browser not to clear these files on its own. As with any download, Hugging Face and jsDelivr see your IP address and which file you fetch, but no content.
What never leaves your device?
- Your audio: files, microphone and shared tab or screen audio. In live mode the audio only stays in memory and is never stored.
- Your results: transcripts, speaker names, summaries and saved sessions are kept only in your browser’s storage (IndexedDB).
- Your exports: TXT, Word, SRT and the other formats are created by your browser as files on your device.
Since none of this is stored with us, we cannot recover it either. Export anything that matters to you.
How can I check this myself?
- Open TranscriptBalance in Chrome, Edge or Firefox and press F12 (or right-click and choose “Inspect”). Switch to the “Network” tab.
- Drop a file and start the transcription.
- You will see downloads from huggingface.co or hf.co and cdn.jsdelivr.net (mostly the first time), parts of the app from our own domain and short messages to
/api/e. There is no request that uploads your file.
Click on an /api/e message to see what it contains: the type of action and rough details such as the model and the length rounded to minutes, never any text or file names.
What leaves your device with optional features?
- Gemini with your own key (live mode only, optional, 18+): your browser then sends excerpts of the transcript and your questions straight to Google, not to us. Without Gemini, live mode uses Chrome’s built-in AI on your device, or no AI at all.
- Chrome AI: Chrome downloads its model from Google itself the first time (about 4 GB). It processes your text on your device.
- Calendar links: clicking “Google Calendar” or “Outlook” sends an event’s title, time and description. The .ics file stays on your device.
- Anonymous usage statistics: we count visits and which features are used, without cookies, without storing IP addresses and never with content. The switch in the app under “About & licenses” stops the feature messages. With Global Privacy Control or Do Not Track enabled, the app sends nothing and your visit is not counted either.
All the details are in the privacy policy.
Requirements and limits
- Browser: WebGPU is available in current Chrome and Edge on a computer, for example, and in other browsers depending on version and device. Without WebGPU everything runs on the processor, except the “Precise” model and summaries with Gemma.
- Speed: with a graphics card, transcription is usually much faster than the recording is long. On weaker devices without WebGPU, a more accurate model can take as long as the recording itself, or longer. The app measures how fast your device is and suggests a model to match.
- Memory: the audio is processed in memory. Very long recordings hit the limits sooner on a phone than on a computer.
- Keep the tab open: the work only happens while the page is open. The app keeps your screen awake and, for jobs with more than 10 minutes of audio, can notify you when it is done.
- Storage space: the models take up space on your device. If your browser clears the site’s data, it downloads them again next time.
Frequently asked questions
Is in-browser transcription GDPR-compliant?
TranscriptBalance never receives your content: audio, transcripts and file names stay on your device, and no server processes them. If you use a recording for more than private purposes, you are responsible for the data protection of the people in it. For interviews or meetings, get their consent and follow the rules of your university or employer.
Can Hugging Face see what I transcribe?
No. Hugging Face only serves the model files and, as with any download, sees your IP address and which model you fetch. For the German live model, that reveals that you are transcribing German. Audio and text never go to Hugging Face.
Why is the first run slower?
The first time, your browser downloads the Whisper model, between 44 and 762 MB depending on the model. After that it is stored in your browser and TranscriptBalance starts without downloading it again.
How do I delete the models and my transcripts?
In TranscriptBalance’s model picker, delete transcription models one by one with “Delete model”; “Delete all” removes every stored model file. Remove transcripts one by one or with “New”, and saved sessions with “Delete session”. To remove everything at once, clear TranscriptBalance’s site data in your browser.
Why does the “Precise” model need a graphics card?
“Precise” is Whisper large-v3-turbo, the largest of TranscriptBalance’s four models. On the processor it would be too slow for long recordings, so the app only offers it with WebGPU. Without WebGPU, “Balanced” (Whisper small) is the most accurate choice.
Does the AI summary run locally too?
In file mode, yes. TranscriptBalance first uses Chrome’s built-in AI (Gemini Nano) and otherwise Gemma 3 1B via WebGPU, both on your device. Only in live mode can you choose to connect Gemini with your own Google key; excerpts of the transcript then go to Google.