OpenAI benchmarks Sol against Opus 5 in comparison charts
1 Sep 22 10:30 AM · 3d ago · 2 articles · 1 source · development 1 of 1
OpenAI published performance comparisons pitting GPT-6 Sol against Opus 5 (not the newer Opus 5.5), showing Sol achieving 33.2% on AutomationBench at 27 cents per task versus Opus 5's 26.9% at 11 times the cost. OpenAI also included a footnote claiming Fable 5.1 fell back to Opus 5 on roughly 40% of tasks. Anthropic countered that Opus 5.5 (which scored 40.0% on the same test) needs fewer tokens to complete jobs.
“Opus 5.5 is our first model since we called for pacing the frontier. As with previous models, it was tested by external evaluators before release, including METR and Frontier Design.”
Anthropic · techmeme ↗Anthropic AI lab releasing Opus 5.5OpenAI AI lab releasing GPT-6 models
Dario Amodei CEO of Anthropic
The whole story articles the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by ZeroHedge News, 3d ago
All 1 developments of Anthropic, OpenAI race to undercut each other on AI model… →
Newswires