Drop in a lecture recording or a downloaded YouTube audio track — up to about two hours — and get back a plain-text transcript. Nothing is uploaded. The speech model downloads once to your browser and every second of audio is transcribed on your own hardware.
Speed depends entirely on your device. With a WebGPU-capable browser (Chrome or Edge, generally) the model runs on your GPU and long files move quickly. Firefox and Safari have limited or no WebGPU support, so the model falls back to CPU (WASM) — a two-hour file with a large model on CPU alone can take a very long time.
Start with Base, Small, or Large v3 Turbo. Only reach for the full Large v3 model if your machine is comfortably keeping up with Turbo.
The model file is cached by your browser after the first download — later visits and transcriptions skip straight to processing.