Developer notes GPT Sol and Astra lead benchmark, awaits Swift evaluation
2 Sep 17 5:27 PM · 8d ago · 2 posts · 2 sources · development 2 of 2
Kelset, a developer on Mastodon, comments that the benchmark results align with their own experience, noting GPT Sol and Astra (along with Fable) rank at the top. They express interest in similar evaluation benchmarks for Swift and SwiftUI development.
“not surprised to see GPT Sol and Astra on top (together with Fable) - matches my experience (still waiting for a good eval for the Swift/SwiftUI side 🫠)”
kelset@mastodon.onlineGoogle Android team Benchmark creator and publisher
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
What people said 1 voice · verbatim
-
K
The Android team released v2 of their benchmark for LLMs: https:// android-developers.googleblog. com/2026/09/android-bench-2-long-horizon-tasks.html not surprised to see GPT Sol and Astra on top (together with Fable) - matches my experience (still waiting for a good eval for the Swift/SwiftUI side 🫠)
All 2 developments of Google releases Android Bench 2.0 with long-horizon AI… →
Google NewsMastodonHacker News