conv.

All stories
AIQuiet 11d · day 11

GoBench benchmark measures LLM reasoning on 9x9 Go against KataGo

A new evaluation framework tests large language models' general reasoning by pitting them against increasingly difficult Go opponents.

What to know

  • GoBench evaluates LLMs on 9x9 Go against KataGo opponents, measuring performance from random to superhuman skill levels.
  • GPT-6 Astra max reaches 2500 Elo while KataGo achieves 4400 Elo, suggesting significant performance gap.
  • The benchmark correlates strongly with ARC-AGI 2 (r=0.83) and remains unsaturated, suggesting room for improvement.
  • Debate centers on whether Go performance reflects general reasoning or arbitrary developer design choices.

The dispute Whether GoBench measures genuine out-of-distribution reasoning ability or reflects intentional training by frontier labs on Go capability. · positions read across 3 posts and comments

some voices

GoBench is a valid general reasoning measure only if frontier labs did not specifically train on Go.

  • “My reasoning was that Go and ARC-AGI are similar in that they are both out of distribution. I assumed that the frontier labs did not do RL on either of those benchmarks.”

    Roland31415 · Reddit ↗
some voices

LLM Go performance is an arbitrary developer choice and does not measure genuine reasoning like ARC-AGI intended to.

  • “In contrast, LLMs doing very well at chess or go is a totally arbitrary decision by the developers. The training data costs $0. It measures general reasoning ability I don't think so.”

    we_are_mammals · Reddit ↗

Roland31415 GoBench creator

GoBench benchmark measures LLM reasoning on 9x9 Go against KataGo
x.com

How it unfolded 4 developments, newest first · click a bar or a number to jump postscomments

Peak 3 pieces in 3h at Sep 16, 2 PM; 4 pieces over 12 days (2 posts · 2 comments) Sep 16, 11 AM — 1 piece · 1 post — X 1Sep 16, 2 PM — 3 pieces · 1 post · 2 comments — Reddit 3Sep 16, 5 PM — quietSep 16, 8 PM — quietSep 16, 11 PM — quietSep 17, 2 AM — quietSep 17, 5 AM — quietSep 17, 8 AM — quietSep 17, 11 AM — quietSep 17, 2 PM — quietSep 17, 5 PM — quietSep 17, 8 PM — quietSep 17, 11 PM — quietSep 18, 2 AM — quietSep 18, 5 AM — quietSep 18, 8 AM — quietSep 18, 11 AM — quietSep 18, 2 PM — quietSep 18, 5 PM — quietSep 18, 8 PM — quietSep 18, 11 PM — quietSep 19, 2 AM — quietSep 19, 5 AM — quietSep 19, 8 AM — quietSep 19, 11 AM — quietSep 19, 2 PM — quietSep 19, 5 PM — quietSep 19, 8 PM — quietSep 19, 11 PM — quietSep 20, 2 AM — quietSep 20, 5 AM — quietSep 20, 8 AM — quietSep 20, 11 AM — quietSep 20, 2 PM — quietSep 20, 5 PM — quietSep 20, 8 PM — quietSep 20, 11 PM — quietSep 21, 2 AM — quietSep 21, 5 AM — quietSep 21, 8 AM — quietSep 21, 11 AM — quietSep 21, 2 PM — quietSep 21, 5 PM — quietSep 21, 8 PM — quietSep 21, 11 PM — quietSep 22, 2 AM — quietSep 22, 5 AM — quietSep 22, 8 AM — quietSep 22, 11 AM — quietSep 22, 2 PM — quietSep 22, 5 PM — quietSep 22, 8 PM — quietSep 22, 11 PM — quietSep 23, 2 AM — quietSep 23, 5 AM — quietSep 23, 8 AM — quietSep 23, 11 AM — quietSep 23, 2 PM — quietSep 23, 5 PM — quietSep 23, 8 PM — quietSep 23, 11 PM — quietSep 24, 2 AM — quietSep 24, 5 AM — quietSep 24, 8 AM — quietSep 24, 11 AM — quietSep 24, 2 PM — quietSep 24, 5 PM — quietSep 24, 8 PM — quietSep 24, 11 PM — quietSep 25, 2 AM — quietSep 25, 5 AM — quietSep 25, 8 AM — quietSep 25, 11 AM — quietSep 25, 2 PM — quietSep 25, 5 PM — quietSep 25, 8 PM — quietSep 25, 11 PM — quietYesterday, 2 AM — quietYesterday, 5 AM — quietYesterday, 8 AM — quietYesterday, 11 AM — quietYesterday, 2 PM — quietYesterday, 5 PM — quietYesterday, 8 PM — quietYesterday, 11 PM — quietToday, 2 AM — quietToday, 5 AM — quietToday, 8 AM — quietToday, 11 AM — quietToday, 2 PM — quietToday, 5 PM — quietToday, 8 PM — quiet 1–4
Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 11:07 PM ET
  1. 4

    Roland31415 responds that Go is out-of-distribution if labs did not specifically train on it

    The creator responds that Go and ARC-AGI are similar as out-of-distribution tasks, and argues GoBench can only measure general reasoning if frontier labs did not specifically train on Go. He notes Go data might exist slightly in pretraining but shouldn't affect results much.

    “It is true that GoBench can only measure general reasoning if the frontier labs do not specifically train on Go.”
    — Roland31415
    • My reasoning was that Go and ARC-AGI are similar in that they are both out of distribution. I assumed that the frontier labs did not do RL on either of those benchmarks. Go data might exist slightly in pretraining, but that shouldnt affect the results much. It is true that GoBench can only measure general reasoning if the frontier labs do not…

      Roland31415r/MachineLearning11d agoview on r/MachineLearning ↗
  2. 3

    Commenter questions whether GoBench measures general reasoning or arbitrary capability

    A commenter challenges the claim that GoBench measures general reasoning, arguing that LLM performance on Go is an arbitrary developer choice unlike ARC-AGI's attempt at true unsolved problems, and that Go training data costs nearly nothing.

    “In contrast, LLMs doing very well at chess or go is a totally arbitrary decision by the developers. The training data costs $0. It measures general reasoning ability I don't think so.”
    — we_are_mammals
    • strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated. This is missing the point of ARC-AGI. They thought that their data would be impossible to generate automatically. Chollet actually claimed something like that in a video interview. They were wrong, but at least they tried. In contrast, LLMs doing very well at…

      we_are_mammalsr/MachineLearning11d agoview on r/MachineLearning ↗
  3. 2

    GoBench details reveal benchmark correlation with ARC-AGI 2 and remaining unsaturation

    Further details show GoBench strongly correlates with ARC-AGI 2 at r=0.83, measures general reasoning ability, and remains unsaturated. With coding tools and two hours of preparation, Codex with Astra achieves 3560 Elo. The creator plans to maintain the leaderboard as long as it remains unsaturated.

    “It measures general reasoning ability, strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated.”
    — Roland31415
  4. 1

    Roland31415 introduces GoBench LLM evaluation framework

    GoBench evaluates large language models by having them play 9x9 Go against a ladder of KataGo opponents from random to superhuman skill. GPT-6 Astra max, the strongest LLM tested, achieves 2500 Elo compared to KataGo's 4400 Elo peak.

    “Introducing GoBench, where LLMs play 9x9 Go against KataGo players, from random to superhuman.”
    — @Roland65821498
    • Introducing GoBench, where LLMs play 9x9 Go against KataGo players, from random to superhuman. The strongest LLM is GPT-6 Astra max at 2500 Elo. GoBench is far from saturation: KataGo can achieve 3300 Elo while costing a million times less, or 4400 Elo while costing 10x less.

      @Roland65821498X11d ago2▲view on X ↗

What people are saying 0 voices from 0 sites · best of 3 · verbatim