Artificial Analysis benchmarks show mixed gains for Sol and Luna
4 Yesterday 3:21 PM · 1d ago · 4 articles · 3 posts · 13 comments · 3 sources · development 4 of 6
Sol improved 2 points on the Coding Agent Index while Luna regressed 2 points; both models cut hallucination rates sharply (Sol from 92% to 60%) but showed regressions on knowledge-work evaluations like GDPval-AA.
“GPT-6 Sol (max) gains 2 points in the Artificial Analysis Coding Agent Index at half the Cost per Task of its predecessor, while GPT-6 Luna (max) regresses by 2 points.”
Artificial AnalysisOpenAI Developer of GPT-6 Sol and LunaAnthropic Developer of Claude Opus 5.5Simon Willison Independent AI commentator/bloggerThariq Shihipar Commentator quoted on Opus 5.5Artificial Analysis Independent AI benchmarking firm
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by Mashable, 1d ago
What people said 13 voices · verbatim
-
I wish there was more transparency on the plus plans usage limits showing actual token usage and prices per model that eats away at remaining usage.Does anyone know how exactly these price differences for example between sol6 and sol5.6 translate to codex percentages? In theory it seems like for "high" on both it should result in ~3x more usage…
-
Simon, love your work, one piece of minor feedback for the individual model pages is to make the font of the model name potentially bigger than (and above) the conversation id (which means nothing to the audience) "2026-09-22T18:28:00 conversation: 01m355zvyw8946qyraa8zpz6h9 id: 01m355zvyx47zxx5c6q6b3fg0m#".I had all the tabs open individually and…
-
From the perspective of “an average person”, ChatGPT is delivering fantastic products.- For general chat and web search, occasional image editing, small coding work, document review etc. ChatGPT Plus is basically limitless and “just works” since 5.6. I’ve yet to give it some task it cannot do.- When given sensible instructions, it hardly annoys…
-
Since I spent my morning fixing a bug in my OpenAI API proxy that completely broke prompt caching and caused my usage limits to burn like kindling, really happy to see some of their new cache tooling:* Prompt caching dashboard: https://platform.openai.com/usage?usage_section=prompt-cachi...* Adjust reasoning effort and tool availability without…
-
Why do they bother creating effort to market all these different models.All I want to know is how old is the model and how much does it cost. I can figure out which one I want to use based on that, assuming that newer models are always better.Trying to convince us there is a difference between GPT-6-Sol and GPT-5.6-Terra or whatnot is ludicrous to…
-
It's surprising but MiMo V2.6 Pro performs better and is cheaper than GPT 6 Sol on my benchmark[1]. Open weight models are really snapping at the heels of the major western models.1 -
-
I just switched from Claude to openAI. I'm surprised at how much easier it is to talk to. Claude always spoke to me with a suspicious side eye as if I was trying to do something naughty. For example I could not get it to help me get an old abandonware game running (sim tower).
-
Many of them still get the layers wrong.They put both legs on the same side of the bike.Even Astra max which actually put one leg on each side of the bike still somehow messed it up because when it added the bike chain, it put the left leg between the bike chain and the frame.
-
Excluding Opus 5.1 from the coding benchmarks is telling. Opus 5 already matches Astra, Opus 5.1 is much better than 5, and 5.5 is much better again.OpenAI seems really competitive in most areas, and extremely competitive on cost, but still behind on coding.
-
I find it very interesting that for both these models we such a clear progression of better images with higher thinking levels from 'hardly useful' to 'pretty nice'. I feel on many other models low and max are much closer.
-
Getting to the point where these headlines depress me. I just wish they would stop getting better. I don't know where my career is gonna be in a few years.
-
Out of all of the benchmarks out there, pelican bicycle bench is the only one I care about. Thank you Simon.
-
This Luna release might potentially be a big deal for computer use automation at scale
All 6 developments of OpenAI and Anthropic trade blows in AI price war: GPT-6… →
MastodonNewswiresXHacker NewsBlueskyGoogle NewsReddit