conv.

All stories
AIQuiet 9d · day 9

Wispr Flow launches Canto, speech model for real-world dictation

New model achieves lowest word error rate on real dictations, outperforming Google, OpenAI, and others in real-time transcription.

What to know

  • Wispr Flow's Canto achieved 3.4% word error rate on real dictations, the lowest among tested models including Google, OpenAI, AssemblyAI, and Deepgram.
  • Evaluation used 10 hours of actual user dictations from over 2,300 speakers across real-world conditions—not read speech datasets—with strict speaker separation to avoid overfitting.
  • On challenging audio conditions (noise, low volume, short utterances), Canto ranked second overall but led all real-time transcription models; Google's Gemini 3.1 Pro was first but is unsuitable for low-latency applications.

Wispr Advanced Interfaces Lab DeveloperGoogle CompetitorOpenAI CompetitorAssemblyAI CompetitorDeepgram Competitor

Wispr Flow launches Canto, speech model for real-world dictation
x.com

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 3 pieces in two hours at Sep 17, 1 PM; 4 pieces over 9 days (1 article · 3 posts) Sep 17, 1 PM — 3 pieces · 1 article · 2 posts — Hacker News 1, Newswires 1, X 1Sep 17, 3 PM — 1 piece · 1 post — Mastodon 1Sep 17, 5 PM — quietSep 17, 7 PM — quietSep 17, 9 PM — quietSep 17, 11 PM — quietSep 18, 1 AM — quietSep 18, 3 AM — quietSep 18, 5 AM — quietSep 18, 7 AM — quietSep 18, 9 AM — quietSep 18, 11 AM — quietSep 18, 1 PM — quietSep 18, 3 PM — quietSep 18, 5 PM — quietSep 18, 7 PM — quietSep 18, 9 PM — quietSep 18, 11 PM — quietSep 19, 1 AM — quietSep 19, 3 AM — quietSep 19, 5 AM — quietSep 19, 7 AM — quietSep 19, 9 AM — quietSep 19, 11 AM — quietSep 19, 1 PM — quietSep 19, 3 PM — quietSep 19, 5 PM — quietSep 19, 7 PM — quietSep 19, 9 PM — quietSep 19, 11 PM — quietSep 20, 1 AM — quietSep 20, 3 AM — quietSep 20, 5 AM — quietSep 20, 7 AM — quietSep 20, 9 AM — quietSep 20, 11 AM — quietSep 20, 1 PM — quietSep 20, 3 PM — quietSep 20, 5 PM — quietSep 20, 7 PM — quietSep 20, 9 PM — quietSep 20, 11 PM — quietSep 21, 1 AM — quietSep 21, 3 AM — quietSep 21, 5 AM — quietSep 21, 7 AM — quietSep 21, 9 AM — quietSep 21, 11 AM — quietSep 21, 1 PM — quietSep 21, 3 PM — quietSep 21, 5 PM — quietSep 21, 7 PM — quietSep 21, 9 PM — quietSep 21, 11 PM — quietSep 22, 1 AM — quietSep 22, 3 AM — quietSep 22, 5 AM — quietSep 22, 7 AM — quietSep 22, 9 AM — quietSep 22, 11 AM — quietSep 22, 1 PM — quietSep 22, 3 PM — quietSep 22, 5 PM — quietSep 22, 7 PM — quietSep 22, 9 PM — quietSep 22, 11 PM — quietSep 23, 1 AM — quietSep 23, 3 AM — quietSep 23, 5 AM — quietSep 23, 7 AM — quietSep 23, 9 AM — quietSep 23, 11 AM — quietSep 23, 1 PM — quietSep 23, 3 PM — quietSep 23, 5 PM — quietSep 23, 7 PM — quietSep 23, 9 PM — quietSep 23, 11 PM — quietSep 24, 1 AM — quietSep 24, 3 AM — quietSep 24, 5 AM — quietSep 24, 7 AM — quietSep 24, 9 AM — quietSep 24, 11 AM — quietSep 24, 1 PM — quietSep 24, 3 PM — quietSep 24, 5 PM — quietSep 24, 7 PM — quietSep 24, 9 PM — quietSep 24, 11 PM — quietYesterday, 1 AM — quietYesterday, 3 AM — quietYesterday, 5 AM — quietYesterday, 7 AM — quietYesterday, 9 AM — quietYesterday, 11 AM — quietYesterday, 1 PM — quietYesterday, 3 PM — quietYesterday, 5 PM — quietYesterday, 7 PM — quietYesterday, 9 PM — quietYesterday, 11 PM — quietToday, 1 AM — quietToday, 3 AM — quietToday, 5 AM — quietToday, 7 AM — quietToday, 9 AM — quietToday, 11 AM — quietToday, 1 PM — quiet 1–2
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 3:36 PM ET
  1. 2

    Wispr publishes detailed Canto evaluation methodology and results

    Wispr Advanced Interfaces Lab released a comprehensive technical post explaining Canto's development and performance. The evaluation compared Canto against Google, OpenAI, AssemblyAI, and Deepgram using a 10-hour dataset of real dictations sampled from Wispr Flow users across applications and use cases. On a 3-hour challenge set with noisy audio, low volume, and short utterances, Canto ranked second behind Gemini 3.1 Pro but led all real-time transcription models tested.

    “Speech recognition models have become remarkably good at transcribing clean audio recorded under controlled conditions. But real dictation rarely happens under those conditions.”
    — Wispr Advanced Interfaces Lab, Developer · source
  2. 1

    Wispr Flow announces Canto speech model

    Wispr Advanced Interfaces Lab introduced Canto, a speech recognition model optimized for real-time dictation in real-world conditions. The model achieved a 3.4% word error rate on evaluation of 9.8 hours of English dictations from over 2,300 speakers.

    “Introducing Canto, our latest speech model for real-time dictation. On our evaluation of 9.8 hours of English Flow dictations from more than 2,300 speakers, Canto achieved a 3.4% word error rate—the lowest among the models we tested.”
    — @WisprFlow
    1. first by HN Frontpage, 9d ago

    • Introducing Canto, our latest speech model for real-time dictation. On our evaluation of 9.8 hours of English Flow dictations from more than 2,300 speakers, Canto achieved a 3.4% word error rate—the lowest among the models we tested. Learn more about Canto from our CSO, Ariya

      @WisprFlowX9d ago97▲view on X ↗