EZ
EZ2Conv

Speech to Text

Turn speech into text two ways: dictate live with your browser's Web Speech API, or upload an audio or video file and transcribe it with a Whisper AI model that runs in your browser. Supports 99+ languages, editing, copy, and download -- free with no login.

Characters
0
Words
0
Sentences
0
Duration
00:00

Transcript

You can edit the text below

How to Use

  1. [Real-time] Select 'Real-time' tab, choose your language, and click 'Start Recording'
  2. [Real-time] Speak clearly into your microphone - text appears in real-time
  3. [File Upload] Select 'File Upload' tab, choose an AI model, and upload an audio/video file
  4. [File Upload] Click 'Transcribe' - the AI model downloads on first use
  5. Edit, copy, or download your transcript when finished

Useful Tips

  • Real-time mode requires Chrome, Edge, or Safari browser
  • File upload works offline after the AI model is downloaded
  • For long files, the 'Tiny' model offers the fastest processing
  • Only File Upload keeps audio on your device - real-time mode sends audio to your browser's speech service unless on-device recognition is available

Two Ways to Turn Speech into Text

What this tool does

Speech to Text offers two independent engines for the same goal. Real-time mode taps the Web Speech API — the SpeechRecognition interface built into modern browsers — and streams words onto the page as you talk: interim guesses first, then a finalized line once you pause. File upload mode instead loads a Whisper AI model through Transformers.js and transcribes a recording you already have, without listening to a microphone at all.

Choosing between the two modes

Reach for real-time dictation when you want to capture thoughts hands-free: meeting notes, a rough draft, or a quick memo you would rather speak than type. It shines during short, live sessions where seeing text appear immediately keeps you in flow.

File upload is the better fit for material that already exists — an interview, a lecture, a voice memo, or the audio track of a video. Once the Whisper model finishes its one-time download, that mode keeps working even with no connection, which suits long recordings and repeated jobs.

A quick walkthrough

For live work, pick your spoken language from the dropdown, press Start Recording, and grant microphone access when the browser asks; your speech fills the transcript and you can edit it afterward. For a recording, switch to File Upload, choose a model — Tiny for speed, Small for stronger multilingual accuracy — and drop in an MP3, WAV, MP4, or similar file. After the model runs, the finished text is ready to copy or download.

Support and things to watch for

Real-time recognition depends on the browser: Chrome, Edge, and Safari implement the Web Speech API, while Firefox does not, so use File Upload there. Recent Chrome versions can install an on-device language model, keeping microphone audio local; otherwise real-time speech is handled by the browser vendor's speech service. The real-time dropdown lists 22 languages, whereas Whisper covers 99+ with automatic detection. File uploads accept common audio and video formats up to 500MB; for video, pulling the audio out first speeds things up and improves the result.

Frequently Asked Questions

Both engines are open with no limits: live dictation through the Web Speech API and file transcription with the Tiny, Base, and Small Whisper models, for audio or video of any length. There is no per-minute metering and no premium models locked behind a paywall.
It depends on the mode. File transcription runs a Whisper model with Transformers.js locally, so the recording and its text never leave your device. Real-time mode uses the browser's Web Speech API, which usually sends microphone audio to the vendor's servers for recognition — unless Chrome has installed an on-device model. EZ2Conv itself never receives your audio.
No account is required. Choose a mode, dictate or upload your recording, then edit, copy, or download the text. There is no email step or personal information to hand over.
Chrome, Edge, and Safari implement the Web Speech API that live mode relies on; Firefox does not include it. In any browser, the File Upload mode still works, because it transcribes with Whisper without depending on that API.
Real-time mode needs the microphone to capture your voice, and the browser asks once. If you decline, that mode cannot listen — but File Upload requires no microphone at all, since it reads an existing recording.
The real-time dropdown offers 22 languages through the Web Speech API, and the file mode with Whisper recognizes 99+ with automatic detection. Select your language before recording to improve accuracy.