Willison's hands-on tests find Opus 5.5 fails its pelican benchmark
5 Sep 22 7:46 PM · 1d ago · 2 articles · 2 posts · 11 comments · 4 sources · development 5 of 6
Simon Willison's informal SVG-pelican test showed Claude Opus 5.5 at max thinking failing to return a result for the first time, while GPT-6 Luna became one of the cheapest models OpenAI has ever shipped and GPT-5.6 Terra lost its remaining reason to exist once priced level with Sol.
“It communicates clearly, it's cheaper per token than Opus 5.0 with the intelligence of Fable 5.1 it's very token efficient and works across every effort level.”
Thariq ShihiparOpenAI Developer of GPT-6 Sol and LunaAnthropic Developer of Claude Opus 5.5Simon Willison Independent AI commentator/bloggerThariq Shihipar Commentator quoted on Opus 5.5Artificial Analysis Independent AI benchmarking firm
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What was reported 2 claims about this development
-
first by Tech Times, 1d ago · also Digit, Tech in Asia, The GitHub Blog, The Neuron, Constellation Research, Wccftech +6
12 more headlines
- OpenAI launches GPT 6 Sol and Luna with Astra-level AI at lower cost, check details Digit · 1d ago
- OpenAI launches GPT-6 Sol, Luna at half the API cost Tech in Asia · 1d ago
- OpenAI's GPT-6 Sol and GPT-6 Luna now available The GitHub Blog · 1d ago
- 😺 GPT-6 Sol / Luna vs. Claude Opus 5.5 LIVE The Neuron · 1d ago
- OpenAI adds GPT-6 Luna and Sol, touts lower prices Constellation Research · 1d ago
- OpenAI Introduces GPT-6 Sol and Luna With 50% Lower API Prices Unite.AI · 1d ago
- OpenAI Unleashes A New Price War, With GPT-6 Sol And GPT-6 Luna Now Priced Below Claude Opus 5.5 And DeepSeek's V4.1 Flash, Respectively, Negating The Rationale For Open-Weight Models Wccftech · 1d ago
- OpenAI launches GPT-6 Sol and Luna with improved intelligence at lower prices Neowin · 1d ago
- OpenAI Launches GPT-6 Sol And Luna With 50% Lower API Pricing Pulse 2.0 · 1d ago
- OpenAI launches GPT-6 Sol and Luna with sharply lower API prices RuntimeWire · 1d ago
- Introducing GPT-6 Sol and Luna OpenAI · 1d ago
- OpenAI Launches GPT-6 Sol and Luna Minutes After Anthropic Drops Claude Opus 5.5 Decrypt · 1d ago
-
first by Unite.AI, 1d ago
What people said 13 voices · verbatim
-
C
📰 #Codex Daily Brief - 2026/09/23 🔥 今日のトップ3 1. GPT-6 Sol / Luna が Codex に投入、API価格は半額 - 複雑コーディングはSol、量産作業はLuna。トークン単価は前世代の半分 - 重要度:★★★★★ - 出典 openai.com/index/introd... techcrunch.com/2026/09/22/o... learn.chatgpt.com/docs/changelog 👇️続く
-
So with GPT-6 Astra, Codex introduced an experimental feature for context management that’s supported to be beneficial for long conversations. Will that experimental feature now also apply to Sol and Luna?https://community.openai.com/t/experimental-context-manageme...I would also like to point out that it was quite predictable that Terra got…
-
Big model release today - I wrote about Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna - plus comparison grids of pelicans by the different model families at different reasoning levels https:// simonwillison.net/2026/Sep/22/ opus-and-sol-and-luna/
-
It seems I'm one of the few Terra users since Astra dropped?When 5.6 dropped I had no weekly limits and I could just drive my work with Sol xhigh and things were great. Once limits were back (and maybe token prices changed iirc) Sol was no longer usable (on Pro or business) unless I was ok with 4 prompts every 5 hours, so I had to switch to Terra…
-
Simon - I believe you've been doing this with a "one-shot" approach. Have you ever considered seeing what the results are with a few more prompts? Maybe 1,2,3 adjustments?Something like the astra MAX is pretty darn good - but something is up with the right wing and the right foot (flipper?)I bet each of these could be modified to be significantly…
-
Here are some somehow standardized pelican tests but for 3d scenes in threejs at threejseval.comLuna 6 High: https://threejseval.com/models/gpt-6-luna-highSol 6 High: https://threejseval.com/models/gpt-6-sol-highYou can compare any other model on the same prompt. Gallery unlocks after 4 votes:
-
Anthropic it is game over when they passed OpenAI at the beginning of the year.Now it is revenge time and OpenAI kind of is trolling Anthropic by simply going into a price war with impressive performance.OpenAI is doing a decent job this year after they recovered. Anthropic needs to offer more payment options and be clear about token usage. The…
-
Interesting that factual error rates have not improved a lot since the last year. Models are getting smarter but not more trustworthy. At this point I would love to have a model that's maybe not as smart but has lower factual inaccuracies and works harder. I.e don't try to find shortcuts as much and does more rigorous self checks.
-
Your benchmark started being gamed by the frontier models a year ago though. The original idea (find a quirky way to test models with something they don't optimize for) is great, but it needs a refresh.
-
What's the relevance of the pelican benchmark when models probably saw it during training? Didn't OpenAI stop testing against SWE-Something because it was tainted?
-
What I overwhelmingly love about that Pelican grid is the two best ones, they've put a neck scarf on to show speed and wind.
-
Luna is really good model that would have been super SOTA at the beginning of the year and its basically free at these prices.
-
Yep. We just switched several classification jobs we run over to gpt-6 luna. Love the cost savings.
All 6 developments of OpenAI and Anthropic trade blows in AI price war: GPT-6… →
MastodonNewswiresXHacker NewsBlueskyGoogle NewsReddit