AI tools · English audio → TXT / SRT
AI audio to text — English
Transcribe a short English WAV or MP3 locally with Whisper. Download a text draft and approximate SRT subtitles without uploading the recording.
Before you start
English speech only; WAV or MP3 up to 10 MiB and 60 seconds. No speaker identification, translation or live recording. Whisper Tiny can miss speech or invent words, especially with noise or silence. Review every transcript.
First use downloads the model plus the inference runtime. Allow extra time and data. A recent desktop browser is recommended; older phones and restricted networks may fail.
Model weights: 41 MB · Xenova/whisper-tiny.en
AI processing runs on your device. When you load a model, your browser contacts jsDelivr and Hugging Face to download software and model files. Those services receive connection information such as your IP address; your selected files and text are not sent to them. Model files may stay in browser cache. Clear site data to remove the cache.
Load the model when you are ready.
How this AI tool works
Your browser decodes the file to mono audio at 16 kHz. Whisper Tiny English predicts text and segment timestamps. TXT and SRT downloads are created locally. Subtitle timing is approximate.
Xenova/whisper-tiny.en · 79fb389fc764e7c395bd330e9531d9d32ada7049 · q8 · Apache-2.0
Model and method sources
Questions about this tool
Can it transcribe Chinese or German?
This first version uses an English-only model. The page has English, German and Chinese interfaces, but the recording must contain English speech.
Are subtitles and names always accurate?
No. This small model can mishear names, numbers, accents or background sounds. Replay the original and edit the draft before publication.
All tools
Link to this tool from your website
Select and copy this HTML into a relevant resource page or README. It links to an empty workspace; no files or results are shared. Attribution is optional.