3Sep 17 12:14 AM · 10d ago · 1 post · 1 source · development 3 of 3
One of the paper's authors publishes a blog post walking through the study's motivation and findings, arguing that despite low raw scores on CritPt (32% for GPT-5.6 Sol and GPT-6 Astra) and Humanity's Last Exam (47%), broken grading may be masking how close frontier models actually are to matching expert physics judgment.
“Apparently, even a model good enough to be declared artificial general intelligence gets roughly half of graduate-level physics wrong.”
The last author also wrote up a quite-readable blogpost here that accompanies the article: https://jsous.github.io/blogs/is-physics-dead/One tidbit I found particularly interesting: "We tested a GPT-based agentic system, which previously succeeded in resolving several open mathematical conjectures, on open problems in theoretical physics. To our…
Starting to feel more and more like chinese room experimentThe models are confidently answering physics questions, treating it as a math problem, but they don't fundamentally "get it" and even recently failed simple "should i drive to car wash" testThe sample efficiency is just crazy lowStill surprising that even with this they managed to saturate…