Free · no signup · no upload

Your audio never leaves your browser.

Transcription runs on your own computer. Drop a file in, watch the text appear, download TXT, SRT or VTT. Nothing is uploaded, nothing is metered.

See how it works
Drop an audio or video file
MP3 · WAV · M4A · FLAC · OGG · MP4 · MOV · WEBM · MKV · AVI
Processed on this device. No length limit.
0 files
uploaded, ever
No limit
on length or minutes
3 formats
TXT · SRT · VTT
Why local

Three things most free transcribers get wrong.

01

They upload your file

Most tools send audio to a server and ask you to trust a policy. Here there is no request to make: the speech model runs inside the browser tab, on your graphics card.

02

They meter the free tier

Thirty minutes a month, then a wall. Local processing costs us nothing per file, so there is no meter, no queue and no account.

03

They lock the export

Plain text, SRT and VTT all download immediately, with timestamps, without an email address.

How it works

Four steps, all of them on your machine.

Step 01

The model downloads once

The first run fetches the speech model, about 200 MB, into your browser cache. Every later file starts instantly.

Step 02

Your file is decoded in a stream

The browser reads the audio a few seconds at a time, so an hour-long meeting recording does not need an hour of memory.

Step 03

The GPU does the listening

WebGPU runs the model on your graphics card, faster than the recording plays back, and the text appears on screen as it goes.

Step 04

You take the text with you

Search inside the transcript, copy it, or download TXT, SRT or VTT. Close the tab and it is gone from this site, because it was never here.

From the blog
All articles
Questions people ask
Is it really free?

Yes. The transcription runs on your own computer, so it costs us nothing per file. The site is paid for by the ads you see around the tool, never by a paid tier or a time limit.

Does my audio get uploaded?

No. Your file is read by the browser on your machine and processed there. Nothing is sent to our servers or anyone else. You can watch the network tab while it runs.

What do I need to run it?

A recent Chrome or Edge on a desktop or laptop with WebGPU support. The page checks before downloading anything and tells you if something is missing. Firefox, Safari and phones are being worked on.

How accurate is it?

It uses the open-source Whisper models from OpenAI. On clear English speech the fast model is good enough for notes and searchable archives; the larger model is close to human accuracy on most recordings. Names and technical terms are the usual weak spot, so read it through.

Is there a file size or length limit?

No hard limit. Audio is decoded in a stream, so a three-hour recording uses about the same memory as a three-minute one. Long files simply take longer.

Which formats can I export?

Plain text, SRT and VTT. SRT works on every video platform and editor; VTT is what HTML video players use; plain text is for notes and search.

Which file types can I drop in?

Audio: MP3, WAV, M4A, AAC, FLAC and OGG. Video: MP4, MOV, WEBM, MKV, and AVI when its sound track is MP3 or PCM. The video itself is ignored; only the audio is read. Windows Media, FLV and DVD formats are not supported yet, and a quick conversion to MP4 in VLC or HandBrake sorts those out.

Which languages are supported?

English for now. Other languages are on the list once the English experience is solid.