Mercury 2.5 LLM debuts at 770 tokens per second with below-average intelligence
New model trades reasoning capability for speed and affordability, sparking debate over whether raw throughput matters when output quality lags.
What to know
- Mercury 2.5 achieves 770 tokens per second throughput but scores at or below median intelligence on benchmarks, raising questions about the utility of extreme speed without reasoning capability.
- Early users report the model underperforms compared to 14B parameter alternatives and is unsuitable even for basic tasks, undermining its cost and speed advantages.
- Commenters identified cheaper open-weight alternatives (DeepSeek v4 Flash, Qwen) available through established providers, questioning Mercury 2.5's competitive positioning.
- Broader debate centers on whether the diffusion LLM architecture underlying Mercury 2.5 has viable use cases, with some arguing frontier labs have abandoned similar approaches.
The dispute Whether Mercury 2.5's speed advantages can overcome its intelligence deficits for any legitimate use case—consensus leans toward no. · positions read across 9 posts and comments
Speed alone doesn't matter if output quality is poor; existing alternatives offer better value.
-
“I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max.”
freakynit · Hacker News ↗
The diffusion LLM approach is fundamentally flawed and has been abandoned by frontier labs.
-
“I honestly think the diffusion LLM approach is a dead end. It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models”
nylonstrung · Hacker News ↗
Beyond a certain throughput threshold (~1000 tps), additional speed offers diminishing returns for human users.
-
“for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps”
pil0u · Hacker News ↗
Mercury 2.5 Language modelArtificial Analysis Benchmarking organization
How it unfolded 5 developments, newest first · click a bar or a number to jump articlespostscomments
-
5
Commenter demonstrates Mercury 2.5 struggles with basic reasoning constraints
A user shared an example showing Mercury 2.5 failed a simple constraint task—writing French without the letter 'e'—and speculated that extreme speed may come at the cost of reasoning capability, questioning whether throughput above 1000 tokens per second matters for human usage.
“I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps…”
— pil0u -
I don't know the model behind this, but it is absurdly bad.> Write me a coherent paragraph in French, without ever using the letter "e".> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."I suppose this is just a demo of how fast an…
-
-
4
Users report Mercury 2.5 underperforms for practical tasks despite speed advantage
Early adopters and commenters shared negative assessments of Mercury 2.5's practical utility. One user stated the model performs on par with 14B models at best, while another noted attempting to use it despite attractive speed and pricing but finding it unsuitable even for basic tasks.
“I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max.”
— freakynit -
Chat Jimmy clocks at 17K tokens per sec burning LLM into the Chip - https://chatjimmy.ai/ - Source:
2 more of the top 3 · 5 posts in this stretch
-
I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max. Even GPT-OSS-20B performs way better than this in my own attempts to use it.I really really wanted to use this because it offers incredible speeds and pricing combinations. But nop.. I still am not using it.. not…
-
At some point the bottleneck becomes tool calling.. and as such, it's preferably if the model is co-hosted (in the same datacenter, at least) with your code repository and all other reference/context it needs (full documentation for most ecosystems, maybe even a copy of common crawl to minimize web fetch usage, etc)
-
-
3
Commenters debate whether speed-optimized LLM architectures have viable use cases
Discussion emerged about the diffusion LLM approach underlying Mercury 2.5, with one commenter arguing it is a 'dead end' and noting that frontier labs like Google have abandoned similar speed-focused experiments, while others questioned what practical applications justify the intelligence trade-offs.
“I honestly think the diffusion LLM approach is a dead end. It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models…”
— nylonstrung -
I honestly think the diffusion LLM approach is a dead endIt's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small modelsStill unclear for what, if any use cases this is pareto frontier
-
-
2
Commenters question Mercury 2.5's competitive position against cheaper alternatives
Multiple commenters compared Mercury 2.5 unfavorably to existing open-weight models like DeepSeek v4 Flash and Qwen, noting that similar-capability models are available at lower cost through established inference providers.
“Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next…so I don't see the point.”
— walrus01 -
Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
1 more of the top 2 · 2 posts in this stretch
-
If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.I've used it on a few for fun projects and its decent but the speed is crazy to watch.
-
-
1
Mercury 2.5 LLM launched with 770 tokens per second throughput
Mercury 2.5 was released with performance metrics showing below-average intelligence (score of 12 on Artificial Analysis Intelligence Index, matching median) but notably fast token generation at 770 tokens per second. The model offers a 260k token context window and moderate pricing at $0.25 per million input tokens and $0.75 per million output tokens.
“Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price. It's also notably fast and fairly concise.”
— Artificial Analysis · source -
1 outlet first by HN Frontpage, 2d ago · read ↗
-
What people are saying 2 voices from 1 site · best of 9 · verbatim
- What are the actual use cases where 770 tokens per second at below-average intelligence is preferable to slower but smarter models?
- At what throughput threshold does additional token generation speed become negligible for human perception?
- Sep 23
-
I'm still sad that we haven't seen a new Taalas style chip a la https://chatjimmy.ai/. Smaller models are good enough now to make that insane burst of tokens so useful.
-
I used this a few days ago and thought something must be wrong with how fast it was responding. "Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price." this is so funny. So when you have a stupid model that is fast - what do you use it for?