Claude Opus 5.5 reaches top spot in Artificial Analysis benchmark
Anthropic's latest model scores highest on industry intelligence index but at elevated pricing.
What to know
- Claude Opus 5.5 tops Artificial Analysis Intelligence Index with a score of 58, more than double the median of 25.
- The model carries premium pricing at $4 per 1M input tokens and $20 per 1M output tokens—both above market median.
- Evaluation showed Opus 5.5 generates significantly more tokens (260M) than comparable models (median 92M), indicating verbose output.
Anthropic AI developerArtificial Analysis Benchmark provider
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Anthropic research account highlights Opus 5.5's benchmark ranking
The MiaAI_lab account on X announces that Opus 5.5 has taken the top spot in Artificial Analysis rankings.
“Opus 5.5 takes the #1 spot in Artificial Analysis…”
— @MiaAI_lab -
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6
2 more of the top 3 · 24 posts in this stretch
-
This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-mediumI've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both…
-
Opus 5.5 takes the #1 spot in Artificial Analysis 😲
-
-
1
Claude Opus 5.5 priced above comparable models
Input tokens cost $4.00 per 1M (median: $2.00) and output tokens cost $20.00 per 1M (median: $10.00), with total evaluation cost of $8,708.20.
“Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price.”
— Artificial Analysis · source -
background
Claude Opus 5.5 rated high in intelligence but verbose in token generation — The model generated 260M tokens during evaluation, significantly higher than the median of 92M tokens, indicating verbose output patterns.
-
background
Artificial Analysis publishes Claude Opus 5.5 benchmark evaluation — Artificial Analysis Intelligence Index v4.3.2 evaluates Claude Opus 5.5 across ten benchmark categories including AA-Briefcase, AutomationBench-AA, Terminal-Bench 4.0, and others, scoring the model at 58—well above the median of 25 for comparable models.
Also covered reported alongside — the timeline has no entry for these yet
-
first by HN Best, 3d ago · also HN Frontpage
1 more headline
- Claude Opus 5.5 Intelligence, Performance and Price Analysis HN Frontpage · 3d ago
What people are saying 21 voices from 1 site · best of 24 · verbatim
- Sep 23
-
Max reasoning seemed to get stuck for me too. My default is Medium which seems to work pretty well both in tight and longer running loops.
-
Pick 10 reasonably intelligent people randomly. How many of them do you think can draw a pelican on a bicycle? I bet most of them can barely even draw a bicycle.This whole test tells me nothing.
-
I am starting to wonder how useful these benchmark still are? Aren't all of these models trained to ace these benchmarks?In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding…
- Sep 22
-
The very first sentence:> Claude Opus 5.5 is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price.What does it mean for a group of similarly priced things to have one that's somewhat more expensive? Cost relative to cost means nothing. You'd think they are would talk about performance…
-
The output style and verbosity with Opus 5.5 is a very big improvement over Opus 5. I predict Opus 5 will be a version with a sudden churn.
-
It sholud definitely not be more capble than Fable. It seems tight comparision but cannot find out its details. I'll use this and check the perceived performance.
-
No surprise then that the default effort level for this model in Claude Code is Medium, even if you had Opus 5 set to High...
-
This continues to show that these foundational models are only slightly better than open weight models but cost around 100x as much. The history of tech is riddled with “good enough” eating “best” for lunch all day long. Unless the big labs come up with a viable business plan pronto it’s looking like AI will be no different.There are no prizes to…
-
Hey! From the Artificial Analysis team. We also have a model releases page which shows all reasoning efforts (not just max), including the trade off curves
-
I'm amazed they didn't test xhigh thinking mode explicitly to ensure it didn't exceed the 128k thinking budget allocation. I guess pace of development gets away from everyone, even OpenAI.
-
Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off…
-
Fingers crossed on this one. I had gone back to 4.8, because 5 was not very good at following instructions or remembering instructions. I found myself repeating quite often what I wanted and what I was trying to do. Opus 5 was more like haiku than it was 4.8 in that respect.
-
"somewhat expensive when comparing to other models of similar price"?That says something about your selected range, and nothing about the model.
-
This is totally a thing I noticed myself about 3 months ago. Medium thinking effort is ideal for most tasks. At high and above, models tend to generate more output in the form of comments or code for the same problem with no real benefit. Its a self-feeding loop: more output becomes more input, which then becomes more output. High is the highest I…
-
"High" to me looks like the one to use. https://artificialanalysis.ai/models/claude-opus-5-5-highMany benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.
-
I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.
-
For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money…
-
"This is a classic test request..."I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.
-
Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they…
-
Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...
-
Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.Edit: