conv.

All stories
AIQuiet 2d · day 3

Mercury 2.5 LLM debuts at 770 tokens per second with below-average intelligence

New model trades reasoning capability for speed and affordability, sparking debate over whether raw throughput matters when output quality lags.

What to know

  • Mercury 2.5 achieves 770 tokens per second throughput but scores at or below median intelligence on benchmarks, raising questions about the utility of extreme speed without reasoning capability.
  • Early users report the model underperforms compared to 14B parameter alternatives and is unsuitable even for basic tasks, undermining its cost and speed advantages.
  • Commenters identified cheaper open-weight alternatives (DeepSeek v4 Flash, Qwen) available through established providers, questioning Mercury 2.5's competitive positioning.
  • Broader debate centers on whether the diffusion LLM architecture underlying Mercury 2.5 has viable use cases, with some arguing frontier labs have abandoned similar approaches.

The dispute Whether Mercury 2.5's speed advantages can overcome its intelligence deficits for any legitimate use case—consensus leans toward no. · positions read across 9 posts and comments

most voices

Speed alone doesn't matter if output quality is poor; existing alternatives offer better value.

  • “I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max.”

    freakynit · Hacker News ↗
some voices

The diffusion LLM approach is fundamentally flawed and has been abandoned by frontier labs.

  • “I honestly think the diffusion LLM approach is a dead end. It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models”

    nylonstrung · Hacker News ↗
some voices

Beyond a certain throughput threshold (~1000 tps), additional speed offers diminishing returns for human users.

  • “for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps”

    pil0u · Hacker News ↗

Mercury 2.5 Language modelArtificial Analysis Benchmarking organization

How it unfolded 5 developments, newest first · click a bar or a number to jump articlespostscomments

Peak 3 pieces in one hour at Sep 23, 10 PM; 12 pieces over 3 days (1 article · 2 posts · 9 comments) Sep 23, 5 PM — 2 pieces · 1 article · 1 post — Hacker News 1, Newswires 1Sep 23, 6 PM — 2 pieces · 2 comments — Hacker News 2Sep 23, 7 PM — 1 piece · 1 comment — Hacker News 1Sep 23, 8 PM — quietSep 23, 9 PM — quietSep 23, 10 PM — 3 pieces · 3 comments — Hacker News 3Sep 23, 11 PM — 2 pieces · 2 comments — Hacker News 2Sep 24, 12 AM — quietSep 24, 1 AM — quietSep 24, 2 AM — quietSep 24, 3 AM — 1 piece · 1 post — Mastodon 1Sep 24, 4 AM — 1 piece · 1 comment — Hacker News 1Sep 24, 5 AM — quietSep 24, 6 AM — quietSep 24, 7 AM — quietSep 24, 8 AM — quietSep 24, 9 AM — quietSep 24, 10 AM — quietSep 24, 11 AM — quietSep 24, 12 PM — quietSep 24, 1 PM — quietSep 24, 2 PM — quietSep 24, 3 PM — quietSep 24, 4 PM — quietSep 24, 5 PM — quietSep 24, 6 PM — quietSep 24, 7 PM — quietSep 24, 8 PM — quietSep 24, 9 PM — quietSep 24, 10 PM — quietSep 24, 11 PM — quietYesterday, 12 AM — quietYesterday, 1 AM — quietYesterday, 2 AM — quietYesterday, 3 AM — quietYesterday, 4 AM — quietYesterday, 5 AM — quietYesterday, 6 AM — quietYesterday, 7 AM — quietYesterday, 8 AM — quietYesterday, 9 AM — quietYesterday, 10 AM — quietYesterday, 11 AM — quietYesterday, 12 PM — quietYesterday, 1 PM — quietYesterday, 2 PM — quietYesterday, 3 PM — quietYesterday, 4 PM — quietYesterday, 5 PM — quietYesterday, 6 PM — quietYesterday, 7 PM — quietYesterday, 8 PM — quietYesterday, 9 PM — quietYesterday, 10 PM — quietYesterday, 11 PM — quietToday, 12 AM — quietToday, 1 AM — quietToday, 2 AM — quietToday, 3 AM — quietToday, 4 AM — quietToday, 5 AM — quietToday, 6 AM — quiet 1–345
Sep 24yesterdaynow · 7:41 AM ET
  1. 5

    Commenter demonstrates Mercury 2.5 struggles with basic reasoning constraints

    A user shared an example showing Mercury 2.5 failed a simple constraint task—writing French without the letter 'e'—and speculated that extreme speed may come at the cost of reasoning capability, questioning whether throughput above 1000 tokens per second matters for human usage.

    “I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps…”
    — pil0u
    • I don't know the model behind this, but it is absurdly bad.> Write me a coherent paragraph in French, without ever using the letter "e".> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."I suppose this is just a demo of how fast an…

      pil0uHacker News2d agoview on Hacker News ↗
  2. 4

    Users report Mercury 2.5 underperforms for practical tasks despite speed advantage

    Early adopters and commenters shared negative assessments of Mercury 2.5's practical utility. One user stated the model performs on par with 14B models at best, while another noted attempting to use it despite attractive speed and pricing but finding it unsuitable even for basic tasks.

    “I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max.”
    — freakynit
    • Chat Jimmy clocks at 17K tokens per sec burning LLM into the Chip - https://chatjimmy.ai/ - Source:

      the_arunHacker News2d agoview on Hacker News ↗
    2 more of the top 3 · 5 posts in this stretch
    • I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max. Even GPT-OSS-20B performs way better than this in my own attempts to use it.I really really wanted to use this because it offers incredible speeds and pricing combinations. But nop.. I still am not using it.. not…

      freakynitHacker News2d agoview on Hacker News ↗
    • At some point the bottleneck becomes tool calling.. and as such, it's preferably if the model is co-hosted (in the same datacenter, at least) with your code repository and all other reference/context it needs (full documentation for most ecosystems, maybe even a copy of common crawl to minimize web fetch usage, etc)

      nextaccounticHacker News2d agoview on Hacker News ↗
    all of them →
  3. 3

    Commenters debate whether speed-optimized LLM architectures have viable use cases

    Discussion emerged about the diffusion LLM approach underlying Mercury 2.5, with one commenter arguing it is a 'dead end' and noting that frontier labs like Google have abandoned similar speed-focused experiments, while others questioned what practical applications justify the intelligence trade-offs.

    “I honestly think the diffusion LLM approach is a dead end. It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models…”
    — nylonstrung
    • I honestly think the diffusion LLM approach is a dead endIt's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small modelsStill unclear for what, if any use cases this is pareto frontier

      nylonstrungHacker News2d agoview on Hacker News ↗
  4. 2

    Commenters question Mercury 2.5's competitive position against cheaper alternatives

    Multiple commenters compared Mercury 2.5 unfavorably to existing open-weight models like DeepSeek v4 Flash and Qwen, noting that similar-capability models are available at lower cost through established inference providers.

    “Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next…so I don't see the point.”
    — walrus01
    • Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.

      walrus01Hacker News2d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.I've used it on a few for fun projects and its decent but the speed is crazy to watch.

      bearjawsHacker News2d agoview on Hacker News ↗
    all of them →
  5. 1

    Mercury 2.5 LLM launched with 770 tokens per second throughput

    Mercury 2.5 was released with performance metrics showing below-average intelligence (score of 12 on Artificial Analysis Intelligence Index, matching median) but notably fast token generation at 770 tokens per second. The model offers a 260k token context window and moderate pricing at $0.25 per million input tokens and $0.75 per million output tokens.

    “Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price. It's also notably fast and fairly concise.”
    — Artificial Analysis · source
    1. 1 outlet first by HN Frontpage, 2d ago · read ↗

What people are saying 2 voices from 1 site · best of 9 · verbatim

Still unanswered
  • What are the actual use cases where 770 tokens per second at below-average intelligence is preferable to slower but smarter models?
  • At what throughput threshold does additional token generation speed become negligible for human perception?