WhisperWeb100% on-device · Hindi + English · private by default
engine…
device —
cache —
model none
Record or drop audio 16 kHz · VAD · auto-segment
00:00.0
idle — tap the dot to record
Transcript
Nothing yet. Load a model on the right, then record or drop an audio file. Everything runs locally — no audio leaves this device.
words 0audio 0s
infer —speed —
1 · Model — English ultrafast
tiny.en fastest · moonshine-tiny 30MB superfast 5-15× vs Whisper · turbo most accurate English (32→4 decoder layers, ~98% large-v3). All cache offline. q8=WASM stable, q4=turbo smallest.
Models fetch once from Hugging Face CDN, then stored in Cache API — offline forever.
2 · Options · English
🇺🇸
ultrafast
English (en) — fixed
No language detection · fastest path
30s
Long files split into overlapping chunks (stride 40%). 30s is Whisper native window.
VAD: silence ≥0.55s closes segment, ≥0.25s opens, 0.25s pre-roll, hard split at 28s.
Extras · On-device AI superfast · WASM
Summarizer 80MB
Sentiment 67MB
Embeddings 22MB
Keywords 0MB
TTS Kokoro 82M
100% on-device
Post-process your transcript locally — no server, all WASM/WebGPU. Open TTS Studio →
New: turbo 809M & moonshine 30MB in Model dropdown above — try them. TTS now has its own page with tiny/small/medium Kokoro models.
All tools stream live per-chunk for long files · Sentiment 3-way, keywords TF-IDF, embeddings cosine.
History local only
Transcripts save here automatically (IndexedDB, never leaves device).
⬇ Drop audio to transcribe (wav · mp3 · flac · m4a · ogg · webm)