conv.

All stories
AIQuiet 5d · day 9

Google releases Android Bench 2.0 with long-horizon AI coding tasks

New benchmark evaluates LLMs and agents on multi-day engineering challenges with continuous scoring instead of binary pass/fail.

What to know

  • Google's Android Bench 2.0 introduces long-horizon tasks simulating multi-day engineering work, with pass rates around 28% versus 91% on simpler original tasks.
  • New continuous scoring system replaces binary pass/fail grading to better reflect partial progress on complex challenges.
  • Analysis shows AI excels at writing new code and deterministic transformations (Java to Kotlin) but struggles with refactoring, migrations, and cross-platform app porting.

Google Android team Benchmark creator and publisher

Google releases Android Bench 2.0 with long-horizon AI coding tasks
android-developers.googleblog.com

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 4 pieces in 3h at Sep 17, 11 AM; 7 pieces over 9 days (4 articles · 3 posts) Sep 17, 11 AM — 4 pieces · 4 articles — Google News 4Sep 17, 2 PM — 1 piece · 1 post — Mastodon 1Sep 17, 5 PM — quietSep 17, 8 PM — quietSep 17, 11 PM — quietSep 18, 2 AM — quietSep 18, 5 AM — quietSep 18, 8 AM — quietSep 18, 11 AM — quietSep 18, 2 PM — quietSep 18, 5 PM — quietSep 18, 8 PM — quietSep 18, 11 PM — quietSep 19, 2 AM — quietSep 19, 5 AM — quietSep 19, 8 AM — quietSep 19, 11 AM — 1 piece · 1 post — Hacker News 1Sep 19, 2 PM — quietSep 19, 5 PM — quietSep 19, 8 PM — quietSep 19, 11 PM — quietSep 20, 2 AM — quietSep 20, 5 AM — quietSep 20, 8 AM — quietSep 20, 11 AM — quietSep 20, 2 PM — quietSep 20, 5 PM — quietSep 20, 8 PM — quietSep 20, 11 PM — quietSep 21, 2 AM — quietSep 21, 5 AM — quietSep 21, 8 AM — quietSep 21, 11 AM — quietSep 21, 2 PM — quietSep 21, 5 PM — quietSep 21, 8 PM — 1 piece · 1 post — Hacker News 1Sep 21, 11 PM — quietSep 22, 2 AM — quietSep 22, 5 AM — quietSep 22, 8 AM — quietSep 22, 11 AM — quietSep 22, 2 PM — quietSep 22, 5 PM — quietSep 22, 8 PM — quietSep 22, 11 PM — quietSep 23, 2 AM — quietSep 23, 5 AM — quietSep 23, 8 AM — quietSep 23, 11 AM — quietSep 23, 2 PM — quietSep 23, 5 PM — quietSep 23, 8 PM — quietSep 23, 11 PM — quietSep 24, 2 AM — quietSep 24, 5 AM — quietSep 24, 8 AM — quietSep 24, 11 AM — quietSep 24, 2 PM — quietSep 24, 5 PM — quietSep 24, 8 PM — quietSep 24, 11 PM — quietYesterday, 2 AM — quietYesterday, 5 AM — quietYesterday, 8 AM — quietYesterday, 11 AM — quietYesterday, 2 PM — quietYesterday, 5 PM — quietYesterday, 8 PM — quietYesterday, 11 PM — quietToday, 2 AM — quietToday, 5 AM — quietToday, 8 AM — quietToday, 11 AM — quietToday, 2 PM — quiet 1–2
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 4:28 PM ET
  1. 2

    Developer notes GPT Sol and Astra lead benchmark, awaits Swift evaluation

    Kelset, a developer on Mastodon, comments that the benchmark results align with their own experience, noting GPT Sol and Astra (along with Fable) rank at the top. They express interest in similar evaluation benchmarks for Swift and SwiftUI development.

    “not surprised to see GPT Sol and Astra on top (together with Fable) - matches my experience (still waiting for a good eval for the Swift/SwiftUI side 🫠)…”
    — kelset@mastodon.online
    • kelset@mastodon.online

      The Android team released v2 of their benchmark for LLMs: https:// android-developers.googleblog. com/2026/09/android-bench-2-long-horizon-tasks.html not surprised to see GPT Sol and Astra on top (together with Fable) - matches my experience (still waiting for a good eval for the Swift/SwiftUI side 🫠)

      kelset@mastodon.onlineMastodon8d agoview on Mastodon ↗
  2. 1

    Android Bench 2.0 reveals significant performance gap on long-horizon tasks

    Results show the highest pass rate for long-horizon tasks is around 28%, dramatically lower than the ~91% pass rate on the original benchmark's simpler tasks. Analysis reveals AI excels at writing new code but struggles with refactoring and migrations due to architectural complexity, and cross-platform app porting remains unsolved by all tested models.

    “The highest pass rate for LHTs is around 28%, much lower than the ~91% for the original tasks in the benchmark.”
    — Google Android team
    1. 2 outlets first by blog.google, 9d ago · also 9to5Google · read ↗

  3. background

    Google shifts from binary to continuous scoring for complex tasks — Google moves away from binary pass/fail grading toward continuous scoring that better captures partial progress on multi-day tasks. Example: an agent that refactors 40 screens to Jetpack Compose, sets up database tables, and passes 90% of requirements but fails one edge-case assertion would score 0% under binary grading but receives a meaningful completion rate under the new system.

  4. background

    Google releases Android Bench 2.0 with long-horizon task evaluation — Google announces Android Bench 2.0, introducing long-horizon tasks (LHTs)—complex multi-day engineering challenges like building apps from scratch, upgrading dependencies, adding features, and converting cross-platform apps to Android. The update also introduces agentic evaluation for testing AI agents from model providers.

Also covered reported alongside — the timeline has no entry for these yet

  1. first by NewsBytes, 9d ago · also Android Headlines

    1 more headline