Commenter demonstrates Mercury 2.5 struggles with basic reasoning constraints
5 Sep 24 5:02 AM · 2d ago · 1 post · 1 comment · 2 sources · development 5 of 5
A user shared an example showing Mercury 2.5 failed a simple constraint task—writing French without the letter 'e'—and speculated that extreme speed may come at the cost of reasoning capability, questioning whether throughput above 1000 tokens per second matters for human usage.
“I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps”
pil0uMercury 2.5 Language modelArtificial Analysis Benchmarking organization
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 1 voice · verbatim
-
I don't know the model behind this, but it is absurdly bad.> Write me a coherent paragraph in French, without ever using the letter "e".> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."I suppose this is just a demo of how fast an…
All 5 developments of Mercury 2.5 LLM debuts at 770 tokens per second with… →
Hacker NewsNewswiresMastodon