AT AI Tools

Speech to Text (Whisper)

Free, private speech-to-text in your browser. Transcribe or translate audio and video with OpenAI Whisper running on your device — record from your mic, then export TXT or SRT subtitles.

🔒 Runs entirely in your browser — nothing is uploaded

Drop an audio or video file here or click to browse

No audio selected.

The first run downloads the Whisper base model (about 77 MB, or about 200 MB with WebGPU) from Hugging Face. It is cached for next time, and your audio never leaves this device.

Advertisement

Whisper speech recognition that runs on your device

This tool turns spoken words into text with Whisper, the open speech recognition model released by OpenAI. Unlike most online transcription services, nothing is sent to a server: the audio is decoded by your browser, resampled to 16 kHz mono and passed to the model in a background Web Worker. The page picks the fastest engine available — WebGPU on graphics cards and browsers that support it, or WebAssembly on the CPU everywhere else. The first run downloads the model weights from Hugging Face; after that they stay in your browser cache, so later sessions start much faster.

You can transcribe interviews, lectures, podcasts, voice memos and meeting recordings, or the audio track of a video file. Short clips can be recorded straight from your microphone. Long recordings are processed in overlapping 30-second windows, and the progress bar shows how many windows are finished.

Transcripts, translations and subtitles

Choose Transcribe to get text in the original language or Translate to English to let Whisper write an English version of speech in another language. Auto-detect listens to the first 30 seconds to identify the language; if your recording starts with music or silence, selecting the language yourself gives better results. Turn on timestamps to see when each sentence was spoken, and download an SRT subtitle file that you can load into video editors, YouTube or media players.

Tips for accurate results

Clear audio matters more than anything else. Record close to the microphone, avoid background music and let one person speak at a time. Whisper base is a compact model, so names, jargon and heavy accents may need a quick manual correction — always proofread before publishing. If your device has little memory, close other heavy tabs before transcribing long files, and keep this tab open until the transcript is complete.

How to use

  1. Add audioDrop an audio or video file onto the box, or press “Record from microphone” and speak.
  2. Choose language and taskLeave the language on Auto-detect or pick it, then choose Transcribe or Translate to English.
  3. Run WhisperPress Transcribe. The model downloads once, then runs on your device while the progress bar fills.
  4. Export the textToggle timestamps if you need them, then copy the transcript or download it as .txt or .srt subtitles.

Frequently asked questions

Is my audio uploaded to a server?
No. The audio is decoded and transcribed inside your browser. The only download is the Whisper model itself, which comes from Hugging Face the first time and is then cached on your device.
How big is the model and how long does it take?
Whisper base is about 77 MB on the WebAssembly (CPU) path and about 200 MB when WebGPU is used. After the first run it loads from the browser cache. On a modern laptop a minute of audio usually takes a few seconds with WebGPU and longer on CPU.
Which languages are supported?
Whisper recognises about 100 languages. The list offers the most common ones, including English, Spanish, German, French, Italian, Portuguese, Dutch, Swedish, Polish, Ukrainian and Russian. Auto-detect guesses the language from the first 30 seconds.
What does “Translate to English” do?
Whisper listens to speech in any supported language and writes the English translation directly, without a separate translation step. Timestamps and SRT export work the same way.
Which file formats work?
Anything your browser can decode: usually MP3, WAV, M4A/AAC, OGG, Opus, FLAC and the audio track of MP4 or WebM videos. If a file fails, convert it to MP3 or WAV and try again.
Advertisement