Two ways to transcribe: live dictation with your microphone using the browser's built-in speech engine, or local transcription of an audio file (MP3, WAV, M4A, OGG, FLAC) using a local neural speech model. Either way your audio never leaves your device.
Idle. Press Start dictation and speak into your microphone.
Live dictation: The live engine is provided by your browser (Chrome/Edge use a cloud recognition service; Firefox/Safari support varies). Chrome and Edge require a secure context (HTTPS) — on localhost it works directly.
What is speech-to-text?
Speech-to-text (speech recognition) converts spoken audio into written words. This page offers two engines: live dictation, which uses the recognition service built into your browser and shows interim results while you speak (they may still change as the engine reconsiders the context before settling on final text), and a local neural model that transcribes an audio file entirely on your device.
Recognition accuracy depends on microphone quality, background noise, speaking pace, accents and domain vocabulary — proofread any automated transcript before you rely on it.