
GPT Transcribe improves on its predecessor but can't catch ElevenLabs, Google, or Mistral on error rates
Quick Answer
OpenAI's new GPT Transcribe and GPT Live Transcribe models improve on error rates, achieving 3.31% but still lag behind ElevenLabs, Google, and Mistral.
Quick Take
Pricing is reduced by 25% to $0.0045 per audio minute, making it competitive but not the leader in the market.
Key Points
- GPT Transcribe processes audio 34 times faster than real-time.
- Word error rate improved by 0.7 percentage points from GPT-4o Transcribe.
- ElevenLabs Scribe v2 leads the market with a 2.3% error rate.
- Mistral's Voxtral Transcribe V2 starts at $0.003 per minute.
- Both models support multiple languages and transcription context.
📖 Reader Mode
~1 min readOpenAI has released GPT Transcribe and GPT Live Transcribe, two new speech recognition models available through its API. GPT Transcribe handles pre-recorded audio files, processing them about 34 times faster than real time. GPT Live Transcribe is built for real-time streaming with low latency.
According to Artificial Analysis, which runs the AA-WER benchmark, GPT Transcribe hits a word error rate of 3.31 percent. That's a 0.7 percentage point improvement over its year-old predecessor GPT-4o Transcribe. Pricing drops 25 percent at the same time, landing at $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and multiple input languages.
In the AA-WER ranking, OpenAI still sits behind several competitors. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro at 2.9 percent and Mistral's Voxtral Small at 3 percent. Mistral recently undercut the market with Voxtral Transcribe V2, starting at just $0.003 per minute.
Full details are in OpenAI's Transcription Guide. The new transcription models complement OpenAI's recently announced Realtime model generation, which also includes the real-time transcription model GPT-Realtime-Whisper.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

