Wispr Flow launches Canto, speech model for real-world dictation
New model achieves lowest word error rate on real dictations, outperforming Google, OpenAI, and others in real-time transcription.
What to know
- Wispr Flow's Canto achieved 3.4% word error rate on real dictations, the lowest among tested models including Google, OpenAI, AssemblyAI, and Deepgram.
- Evaluation used 10 hours of actual user dictations from over 2,300 speakers across real-world conditions—not read speech datasets—with strict speaker separation to avoid overfitting.
- On challenging audio conditions (noise, low volume, short utterances), Canto ranked second overall but led all real-time transcription models; Google's Gemini 3.1 Pro was first but is unsuitable for low-latency applications.
Wispr Advanced Interfaces Lab DeveloperGoogle CompetitorOpenAI CompetitorAssemblyAI CompetitorDeepgram Competitor
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Wispr publishes detailed Canto evaluation methodology and results
Wispr Advanced Interfaces Lab released a comprehensive technical post explaining Canto's development and performance. The evaluation compared Canto against Google, OpenAI, AssemblyAI, and Deepgram using a 10-hour dataset of real dictations sampled from Wispr Flow users across applications and use cases. On a 3-hour challenge set with noisy audio, low volume, and short utterances, Canto ranked second behind Gemini 3.1 Pro but led all real-time transcription models tested.
“Speech recognition models have become remarkably good at transcribing clean audio recorded under controlled conditions. But real dictation rarely happens under those conditions.”
— Wispr Advanced Interfaces Lab, Developer · source -
1
Wispr Flow announces Canto speech model
Wispr Advanced Interfaces Lab introduced Canto, a speech recognition model optimized for real-time dictation in real-world conditions. The model achieved a 3.4% word error rate on evaluation of 9.8 hours of English dictations from over 2,300 speakers.
“Introducing Canto, our latest speech model for real-time dictation. On our evaluation of 9.8 hours of English Flow dictations from more than 2,300 speakers, Canto achieved a 3.4% word error rate—the lowest among the models we tested.”
— @WisprFlow -
first by HN Frontpage, 9d ago
-
Introducing Canto, our latest speech model for real-time dictation. On our evaluation of 9.8 hours of English Flow dictations from more than 2,300 speakers, Canto achieved a 3.4% word error rate—the lowest among the models we tested. Learn more about Canto from our CSO, Ariya
-