conv.

All stories
AIQuiet 2d · day 4

Claude Opus 5.5 reaches top spot in Artificial Analysis benchmark

Anthropic's latest model scores highest on industry intelligence index but at elevated pricing.

What to know

  • Claude Opus 5.5 tops Artificial Analysis Intelligence Index with a score of 58, more than double the median of 25.
  • The model carries premium pricing at $4 per 1M input tokens and $20 per 1M output tokens—both above market median.
  • Evaluation showed Opus 5.5 generates significantly more tokens (260M) than comparable models (median 92M), indicating verbose output.

Anthropic AI developerArtificial Analysis Benchmark provider

Claude Opus 5.5 reaches top spot in Artificial Analysis benchmark
x.com

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 8 pieces in one hour at Sep 22, 1 PM; 31 pieces over 4 days (2 articles · 7 posts · 22 comments) Sep 22, 12 PM — 4 pieces · 2 articles · 2 posts — Newswires 2, Hacker News 1, X 1Sep 22, 1 PM — 8 pieces · 8 comments — Hacker News 8Sep 22, 2 PM — 2 pieces · 1 post · 1 comment — Hacker News 2Sep 22, 3 PM — 4 pieces · 4 comments — Hacker News 4Sep 22, 4 PM — 1 piece · 1 comment — Hacker News 1Sep 22, 5 PM — quietSep 22, 6 PM — 2 pieces · 2 posts — Mastodon 1, Reddit 1Sep 22, 7 PM — 1 piece · 1 comment — Hacker News 1Sep 22, 8 PM — 3 pieces · 3 comments — Hacker News 3Sep 22, 9 PM — quietSep 22, 10 PM — 1 piece · 1 comment — Hacker News 1Sep 22, 11 PM — 1 piece · 1 post — Mastodon 1Sep 23, 12 AM — quietSep 23, 1 AM — quietSep 23, 2 AM — quietSep 23, 3 AM — 1 piece · 1 comment — Hacker News 1Sep 23, 4 AM — quietSep 23, 5 AM — 1 piece · 1 comment — Hacker News 1Sep 23, 6 AM — quietSep 23, 7 AM — quietSep 23, 8 AM — quietSep 23, 9 AM — quietSep 23, 10 AM — quietSep 23, 11 AM — quietSep 23, 12 PM — quietSep 23, 1 PM — 1 piece · 1 comment — Hacker News 1Sep 23, 2 PM — quietSep 23, 3 PM — quietSep 23, 4 PM — quietSep 23, 5 PM — quietSep 23, 6 PM — quietSep 23, 7 PM — quietSep 23, 8 PM — quietSep 23, 9 PM — 1 piece · 1 post — X 1Sep 23, 10 PM — quietSep 23, 11 PM — quietSep 24, 12 AM — quietSep 24, 1 AM — quietSep 24, 2 AM — quietSep 24, 3 AM — quietSep 24, 4 AM — quietSep 24, 5 AM — quietSep 24, 6 AM — quietSep 24, 7 AM — quietSep 24, 8 AM — quietSep 24, 9 AM — quietSep 24, 10 AM — quietSep 24, 11 AM — quietSep 24, 12 PM — quietSep 24, 1 PM — quietSep 24, 2 PM — quietSep 24, 3 PM — quietSep 24, 4 PM — quietSep 24, 5 PM — quietSep 24, 6 PM — quietSep 24, 7 PM — quietSep 24, 8 PM — quietSep 24, 9 PM — quietSep 24, 10 PM — quietSep 24, 11 PM — quietYesterday, 12 AM — quietYesterday, 1 AM — quietYesterday, 2 AM — quietYesterday, 3 AM — quietYesterday, 4 AM — quietYesterday, 5 AM — quietYesterday, 6 AM — quietYesterday, 7 AM — quietYesterday, 8 AM — quietYesterday, 9 AM — quietYesterday, 10 AM — quietYesterday, 11 AM — quietYesterday, 12 PM — quietYesterday, 1 PM — quietYesterday, 2 PM — quietYesterday, 3 PM — quietYesterday, 4 PM — quietYesterday, 5 PM — quietYesterday, 6 PM — quietYesterday, 7 PM — quietYesterday, 8 PM — quietYesterday, 9 PM — quietYesterday, 10 PM — quietYesterday, 11 PM — quietToday, 12 AM — quietToday, 1 AM — quietToday, 2 AM — quietToday, 3 AM — quietToday, 4 AM — quietToday, 5 AM — quietToday, 6 AM — quietToday, 7 AM — quietToday, 8 AM — quiet 1–2
Sep 23Sep 24yesterdaynow · 9:20 AM ET
  1. 2

    Anthropic research account highlights Opus 5.5's benchmark ranking

    The MiaAI_lab account on X announces that Opus 5.5 has taken the top spot in Artificial Analysis rankings.

    “Opus 5.5 takes the #1 spot in Artificial Analysis…”
    — @MiaAI_lab
    • Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6

      @ArtificialAnlysX2d ago541▲view on X ↗
    2 more of the top 3 · 24 posts in this stretch
    • This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-mediumI've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both…

      simonwHacker News3d agoview on Hacker News ↗
    • Opus 5.5 takes the #1 spot in Artificial Analysis 😲

      @MiaAI_labX3d ago36▲view on X ↗
    all of them →
  2. 1

    Claude Opus 5.5 priced above comparable models

    Input tokens cost $4.00 per 1M (median: $2.00) and output tokens cost $20.00 per 1M (median: $10.00), with total evaluation cost of $8,708.20.

    “Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price.”
    — Artificial Analysis · source
  3. background

    Claude Opus 5.5 rated high in intelligence but verbose in token generation — The model generated 260M tokens during evaluation, significantly higher than the median of 92M tokens, indicating verbose output patterns.

  4. background

    Artificial Analysis publishes Claude Opus 5.5 benchmark evaluation — Artificial Analysis Intelligence Index v4.3.2 evaluates Claude Opus 5.5 across ten benchmark categories including AA-Briefcase, AutomationBench-AA, Terminal-Bench 4.0, and others, scoring the model at 58—well above the median of 25 for comparable models.

Also covered reported alongside — the timeline has no entry for these yet

  1. first by HN Best, 3d ago · also HN Frontpage

    1 more headline

What people are saying 21 voices from 1 site · best of 24 · verbatim