Anthropic research account highlights Opus 5.5's benchmark ranking
2 Sep 22 1:14 PM · 3d ago · 1 post · 1 source · development 2 of 2
The MiaAI_lab account on X announces that Opus 5.5 has taken the top spot in Artificial Analysis rankings.
“Opus 5.5 takes the #1 spot in Artificial Analysis”
@MiaAI_labAnthropic AI developerArtificial Analysis Benchmark provider
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
What people said 24 voices · verbatim
-
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6
-
This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-mediumI've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both…
-
Opus 5.5 takes the #1 spot in Artificial Analysis 😲
-
I am starting to wonder how useful these benchmark still are? Aren't all of these models trained to ace these benchmarks?In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding…
-
This is totally a thing I noticed myself about 3 months ago. Medium thinking effort is ideal for most tasks. At high and above, models tend to generate more output in the form of comments or code for the same problem with no real benefit. Its a self-feeding loop: more output becomes more input, which then becomes more output. High is the highest I…
-
For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money…
-
This continues to show that these foundational models are only slightly better than open weight models but cost around 100x as much. The history of tech is riddled with “good enough” eating “best” for lunch all day long. Unless the big labs come up with a viable business plan pronto it’s looking like AI will be no different.There are no prizes to…
-
Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off…
-
The very first sentence:> Claude Opus 5.5 is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price.What does it mean for a group of similarly priced things to have one that's somewhat more expensive? Cost relative to cost means nothing. You'd think they are would talk about performance…
-
"High" to me looks like the one to use. https://artificialanalysis.ai/models/claude-opus-5-5-highMany benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.
-
Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.Edit:
-
Hey! From the Artificial Analysis team. We also have a model releases page which shows all reasoning efforts (not just max), including the trade off curves
-
Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they…
-
Fingers crossed on this one. I had gone back to 4.8, because 5 was not very good at following instructions or remembering instructions. I found myself repeating quite often what I wanted and what I was trying to do. Opus 5 was more like haiku than it was 4.8 in that respect.
-
Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...
-
"This is a classic test request..."I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
-
I'm amazed they didn't test xhigh thinking mode explicitly to ensure it didn't exceed the 128k thinking budget allocation. I guess pace of development gets away from everyone, even OpenAI.
-
Pick 10 reasonably intelligent people randomly. How many of them do you think can draw a pelican on a bicycle? I bet most of them can barely even draw a bicycle.This whole test tells me nothing.
-
I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.
-
It sholud definitely not be more capble than Fable. It seems tight comparision but cannot find out its details. I'll use this and check the perceived performance.
-
"somewhat expensive when comparing to other models of similar price"?That says something about your selected range, and nothing about the model.
-
The output style and verbosity with Opus 5.5 is a very big improvement over Opus 5. I predict Opus 5 will be a version with a sudden churn.
-
Max reasoning seemed to get stuck for me too. My default is Medium which seems to work pretty well both in tight and longer running loops.
-
No surprise then that the default effort level for this model in Claude Code is Medium, even if you had Opus 5 set to High...
All 2 developments of Claude Opus 5.5 reaches top spot in Artificial Analysis… →
NewswiresXHacker NewsMastodonReddit