conv.

All stories
AIQuiet 7d · day 8

Brood War Bench pits AI coding agents against each other in StarCraft

A hobbyist built a benchmark that lets Codex, Grok, Claude Fable and other models play full games of Brood War, exposing sharply different failure modes.

What to know

  • Brood War Bench lets LLM agents (Codex, Grok 4.6, Claude Fable, others) play full StarCraft: Brood War games against each other with no human control.
  • Models show distinct failure patterns: Codex's harassment tactics work but its production and coordination are weak; Grok reasons extensively but acts rarely; Claude Fable is the most ambitious economically.
  • Even the best-performing agents remain far below basic human skill — the creator notes a beginner doing a simple photon rush would beat every model tested.
  • Commenters situate the project against older AI-StarCraft research (a 2010 UCSC Brood War AI tournament, DeepMind's SC2 work) and a separate LLM Go benchmark (GoBench), and ask for clarification on how agents actually interface with the game.

The dispute Whether Brood War Bench meaningfully differentiates model capability, versus commenters pointing to alternatives like GoBench as showing clearer capability gaps. · positions read across 40 posts and comments

many voices

The post mainly triggers personal nostalgia about StarCraft/Brood War as a formative community and hobby, largely separate from the AI angle.

some voices

The benchmark's methodology and real capability differentiation are unclear, and other approaches (older AI tournaments, GoBench) are cited for comparison.

  • “Did it play by looking at screenshots and sending clicks, or was there other mediation/symbolization?”

    gadtfly · Hacker News ↗
some voices

StarCraft's three factions offer a useful playful metaphor for how to think about deploying different kinds of AI agents.

  • “Protoss: powerful and expensive frontier coding agents you directly micromanage for the toughest tasks”

    mcteamster · Hacker News ↗

“This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.”

benswerd, Benchmark creator · bw.swerdlow.dev/report ↗ · Sep 18

benswerd Creator of Brood War BenchOpenAI Codex AI agent tested in the benchmarkGrok 4.6 (xAI) AI agent tested in the benchmarkClaude Fable (Anthropic) AI agent tested in the benchmark

Brood War Bench pits AI coding agents against each other in StarCraft
bw.swerdlow.dev

How it unfolded 3 developments, newest first · click a bar or a number to jump articlespostscomments

Peak 10 pieces in two hours at Sep 19, 4 PM; 43 pieces over 8 days (2 articles · 6 posts · 35 comments) Sep 19, 10 AM — 3 pieces · 2 articles · 1 post — Newswires 2, Hacker News 1Sep 19, 12 PM — quietSep 19, 2 PM — quietSep 19, 4 PM — 10 pieces · 10 comments — Hacker News 10Sep 19, 6 PM — 2 pieces · 1 post · 1 comment — Hacker News 1, Mastodon 1Sep 19, 8 PM — 3 pieces · 3 comments — Hacker News 3Sep 19, 10 PM — 2 pieces · 2 comments — Hacker News 2Sep 20, 12 AM — 1 piece · 1 post — Mastodon 1Sep 20, 2 AM — 6 pieces · 6 comments — Hacker News 6Sep 20, 4 AM — 6 pieces · 2 posts · 4 comments — Hacker News 4, Mastodon 2Sep 20, 6 AM — 3 pieces · 1 post · 2 comments — Hacker News 2, Mastodon 1Sep 20, 8 AM — 2 pieces · 2 comments — Hacker News 2Sep 20, 10 AM — 3 pieces · 3 comments — Hacker News 3Sep 20, 12 PM — 1 piece · 1 comment — Hacker News 1Sep 20, 2 PM — 1 piece · 1 comment — Hacker News 1Sep 20, 4 PM — quietSep 20, 6 PM — quietSep 20, 8 PM — quietSep 20, 10 PM — quietSep 21, 12 AM — quietSep 21, 2 AM — quietSep 21, 4 AM — quietSep 21, 6 AM — quietSep 21, 8 AM — quietSep 21, 10 AM — quietSep 21, 12 PM — quietSep 21, 2 PM — quietSep 21, 4 PM — quietSep 21, 6 PM — quietSep 21, 8 PM — quietSep 21, 10 PM — quietSep 22, 12 AM — quietSep 22, 2 AM — quietSep 22, 4 AM — quietSep 22, 6 AM — quietSep 22, 8 AM — quietSep 22, 10 AM — quietSep 22, 12 PM — quietSep 22, 2 PM — quietSep 22, 4 PM — quietSep 22, 6 PM — quietSep 22, 8 PM — quietSep 22, 10 PM — quietSep 23, 12 AM — quietSep 23, 2 AM — quietSep 23, 4 AM — quietSep 23, 6 AM — quietSep 23, 8 AM — quietSep 23, 10 AM — quietSep 23, 12 PM — quietSep 23, 2 PM — quietSep 23, 4 PM — quietSep 23, 6 PM — quietSep 23, 8 PM — quietSep 23, 10 PM — quietSep 24, 12 AM — quietSep 24, 2 AM — quietSep 24, 4 AM — quietSep 24, 6 AM — quietSep 24, 8 AM — quietSep 24, 10 AM — quietSep 24, 12 PM — quietSep 24, 2 PM — quietSep 24, 4 PM — quietSep 24, 6 PM — quietSep 24, 8 PM — quietSep 24, 10 PM — quietSep 25, 12 AM — quietSep 25, 2 AM — quietSep 25, 4 AM — quietSep 25, 6 AM — quietSep 25, 8 AM — quietSep 25, 10 AM — quietSep 25, 12 PM — quietSep 25, 2 PM — quietSep 25, 4 PM — quietSep 25, 6 PM — quietSep 25, 8 PM — quietSep 25, 10 PM — quietYesterday, 12 AM — quietYesterday, 2 AM — quietYesterday, 4 AM — quietYesterday, 6 AM — quietYesterday, 8 AM — quietYesterday, 10 AM — quietYesterday, 12 PM — quietYesterday, 2 PM — quietYesterday, 4 PM — quietYesterday, 6 PM — quietYesterday, 8 PM — quietYesterday, 10 PM — quietToday, 12 AM — quietToday, 2 AM — quietToday, 4 AM — quietToday, 6 AM — quietToday, 8 AM — quietToday, 10 AM — quietToday, 12 PM — quiet 1–3
Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 2:15 PM ET
  1. 3

    Report is mirrored and picked up beyond Hacker News

    The same report text was republished/linked via a Mastodon bot account tracking HN frontpage stories, spreading the findings to a wider audience.

    “Protoss: powerful and expensive frontier coding agents you directly micromanage for the toughest tasks…”
    — mcteamster, HN commenter · source
    • greg76@mas.to

      Have trouble picking the right # benchmark about AI agents/models? also played # rts # games ~30 years ago? then you might find Brood War Bench fascinating 😜 https:// bw.

      greg76@mas.toMastodon7d ago1▲view on Mastodon ↗
    2 more of the top 3 · 31 posts in this stretch
    • This is a very interesting benchmark, and I think it has a lot of potential to make the speed of a model quantifiable.I'm often asking myself is it better to use higher or lower effort levels, or to maybe drop down to a "dumber" but faster model. And so using a real-time based competition as a benchmark could shed some light on this, I think.In…

      alembic_fumesHacker News7d agoview on Hacker News ↗
    • hn250@social.lansky.name

      Brood War Bench Link: https:// bw.swerdlow.dev/report Discussion: https:// news.ycombinator.com/item?id=4 9766966

      hn250@social.lansky.nameMastodon7d agoview on Mastodon ↗
    all of them →
  2. 2

    HN commenters question methodology and compare it to prior AI-StarCraft work

    Commenters asked whether agents played in real time via screenshots/clicks or through some other symbolic interface, and drew comparisons to a 2010 UC Santa Cruz Brood War AI tournament and to DeepMind's SC2 agents, as well as a separate LLM benchmark on 9x9 Go (GoBench).

    “Did it play by looking at screenshots and sending clicks, or was there other mediation/symbolization?”
    — gadtfly
    1. 2 outlets Brood War Bench

      first by HN Best, 8d ago · also HN Frontpage

    • Back in 2010, during the early days of bwapi, there was a Brood War AI tournament held by the Expressive Intelligence Studio at UC Santa Cruz. It's interesting to see how different the approaches were back then, vs this or Deepmind's SC2 work.https://web.archive.org/web/20091124210529/http://eis.ucsc.e...There's a great contemporary Ars Technica…

      AntiRushHacker News7d agoview on Hacker News ↗
    2 more of the top 3 · 9 posts in this stretch
    • Unrelated to the benchmark...I love StarCraft. I started playing it right from the beginning, most of my friends right now are from that era. I literally met people that have spread to almost every continent when I was in my early teens. We played at internet cafes and did not have access to the internet, that was priced differently...I miss those…

      pelagicAustralHacker News7d agoview on Hacker News ↗
    • A friend of mine created GoBench[1][2] that evaluates LLMs on 9×9 Go using KataGo opponents as Elo anchors, you see real capability differences there, like Astra Max substantially leading all other models. I think strategy is a generally interesting area to evaluate LLMs on[1]

      GodelNumberingHacker News7d agoview on Hacker News ↗
    all of them →
  3. 1

    Report details model-specific strategies and failures

    The write-up describes Codex's effective but shallow Probe-harassment tactics, its subagents' poor coordination, Grok 4.6's long reasoning traces with almost no issued commands, and Claude Fable's comparatively ambitious tech-climbing and economy-building, while stressing that a beginner human doing a photon rush would beat every model tested.

    “Codex's strongest recurring idea was disruption.”
    — benswerd
  4. background

    benswerd publishes Brood War Bench, an AI-agent StarCraft benchmark — The project grew out of an agent-only version of Brood War the creator built to play with friends; after friends won by simply telling their agent to attack, benswerd built a formal benchmark testing how far AI agents can get on their own.

What people are saying 17 voices from 2 sites · best of 40 · verbatim

Still unanswered
  • Did the agents play in true real time via screen input, or through some discretized/turn-based interface?