Yale-led audit finds physics AI benchmarks are broken, not just hard
Expert re-grading of CritPt and Humanity's Last Exam suggests frontier models are far closer to saturating physics benchmarks than raw scores show.
What to know
- Raw leaderboard scores show frontier models struggling with physics: 32% on CritPt (GPT-5.6 Sol, unimproved by GPT-6 Astra) and 47% on the physics portion of Humanity's Last Exam.
- Yale physics faculty and graduate researchers manually re-graded model answers and found the automated evaluation itself was broken, not just the models' performance weak.
- The study's title claims 'near-saturation of leading benchmarks' once broken grading is corrected, suggesting physics may be less of a holdout against AI capability than leaderboards imply.
- The paper is gaining wider attention after moving from a low-traffic arXiv submission to the Hacker News front page within days.
jsous Paper co-author and blog post authorArman Cohan Co-author of the benchmark re-grading studyYale physics faculty and graduate researchers Conducted expert audits of benchmark gradingOpenAI Developer of the evaluated frontier models
How it unfolded 3 developments, newest first · click a bar or a number to jump articlesposts
-
3
Co-author jsous publishes 'Is Physics Dead' explainer
One of the paper's authors publishes a blog post walking through the study's motivation and findings, arguing that despite low raw scores on CritPt (32% for GPT-5.6 Sol and GPT-6 Astra) and Humanity's Last Exam (47%), broken grading may be masking how close frontier models actually are to matching expert physics judgment.
“Apparently, even a model good enough to be declared artificial general intelligence gets roughly half of graduate-level physics wrong.”
— jsous -
The last author also wrote up a quite-readable blogpost here that accompanies the article: https://jsous.github.io/blogs/is-physics-dead/One tidbit I found particularly interesting: "We tested a GPT-based agentic system, which previously succeeded in resolving several open mathematical conjectures, on open problems in theoretical physics. To our…
2 more of the top 3 · 3 posts in this stretch
-
Starting to feel more and more like chinese room experimentThe models are confidently answering physics questions, treating it as a math problem, but they don't fundamentally "get it" and even recently failed simple "should i drive to car wash" testThe sample efficiency is just crazy lowStill surprising that even with this they managed to saturate…
-
Are models able to do math now, or do they still rely on “tools” to do the math?
-
-
2
Paper gains traction on Hacker News front page
The same arXiv paper is resubmitted and climbs to the Hacker News front page, drawing 85 points and 43 comments as the finding spreads beyond the initial small-scale post.
-
(Trained physicist here)From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced…
2 more of the top 3 · 4 posts in this stretch
-
Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks"John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.When hand grading instead, they found out that the…
-
This is interesting and actually very important for robotics.I've been waiting for this, but all companies seem to not care much now.There is a way out of this by supplying right context (needs a bit of expertise in physics)1 more year and frontier will become crazy good at this as well.
-
- 1 day quiet
-
1
Researchers post physics benchmark re-grading study on arXiv
A paper titled 'How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks' appears on arXiv, reporting that Yale physics faculty and graduate researchers manually audited how frontier models were scored on physics benchmarks.
-
first by HN Frontpage, 11d ago
-
What people are saying 1 voices from 1 site · best of 7 · verbatim
- Sep 16
-
I would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.