Commenter questions whether GoBench measures general reasoning or arbitrary capability
3 Sep 16 4:33 PM · 11d ago · 1 comment · 1 source · development 3 of 4
A commenter challenges the claim that GoBench measures general reasoning, arguing that LLM performance on Go is an arbitrary developer choice unlike ARC-AGI's attempt at true unsolved problems, and that Go training data costs nearly nothing.
“In contrast, LLMs doing very well at chess or go is a totally arbitrary decision by the developers. The training data costs $0. It measures general reasoning ability I don't think so.”
we_are_mammalsRoland31415 GoBench creator
The whole story postscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 1 voice · verbatim
-
strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated. This is missing the point of ARC-AGI. They thought that their data would be impossible to generate automatically. Chollet actually claimed something like that in a video interview. They were wrong, but at least they tried. In contrast, LLMs doing very well at…
All 4 developments of GoBench benchmark measures LLM reasoning on 9x9 Go against… →
RedditX