Speech to text, on your device
Transcribe anything, no minute limits
Transcription services meter you by the minute and want your recording on their servers. This one runs OpenAI's Whisper model inside your browser tab, so interviews, lectures and client calls are transcribed without ever leaving your device — and without a meter.
What you get from local transcription
Nothing uploaded
Confidential interviews and client calls stay on your machine.
Subtitles included
Export .srt with timestamps, ready for any video editor.
Top searched languages
English, Hindi, Spanish, French, Arabic, Bengali and other high-demand languages stay one tap away.
How to transcribe audio to text
- 1
Open the transcriber and pick your file
Drop an audio or video file — MP3, WAV, M4A, OGG, MP4 or WebM — onto the panel above, or use the file picker. Nothing is uploaded; the file is read straight from your device.
- 2
Choose the language and task
Leave it on auto-detect for most recordings, or pick the spoken language for better accuracy. Switch the task to 'Translate to English' if you want an English transcript of non-English speech.
- 3
Start the transcription
The Whisper speech model downloads once and is cached by your browser, then runs locally in 30-second chunks so long recordings progress steadily. There is no minute limit and no account.
- 4
Copy or export the text
Copy the plain transcript, or export a timestamped .srt subtitle file that drops straight into any video editor.
Common questions
Why is there a download the first time?
The speech recognition model itself has to reach your device before it can run. It is fetched once, cached by your browser, and reused instantly on every later visit.
How long does a recording take?
Roughly real-time to a few times faster on a modern laptop with the fast model. Long recordings are processed in 30-second chunks, so progress is steady rather than all-at-once.
Which files can I transcribe?
Any audio or video your browser can decode — MP3, WAV, M4A, OGG, MP4, WebM and more. The audio track is extracted and resampled automatically.
How accurate is it?
Whisper is strong on clear speech and holds up well with accents. Heavy background noise, crosstalk and very quiet recordings are where any model, local or cloud, starts to struggle.