conv.

All stories
AIQuiet 11d · day 14

Yale-led audit finds physics AI benchmarks are broken, not just hard

Expert re-grading of CritPt and Humanity's Last Exam suggests frontier models are far closer to saturating physics benchmarks than raw scores show.

What to know

  • Raw leaderboard scores show frontier models struggling with physics: 32% on CritPt (GPT-5.6 Sol, unimproved by GPT-6 Astra) and 47% on the physics portion of Humanity's Last Exam.
  • Yale physics faculty and graduate researchers manually re-graded model answers and found the automated evaluation itself was broken, not just the models' performance weak.
  • The study's title claims 'near-saturation of leading benchmarks' once broken grading is corrected, suggesting physics may be less of a holdout against AI capability than leaderboards imply.
  • The paper is gaining wider attention after moving from a low-traffic arXiv submission to the Hacker News front page within days.

jsous Paper co-author and blog post authorArman Cohan Co-author of the benchmark re-grading studyYale physics faculty and graduate researchers Conducted expert audits of benchmark gradingOpenAI Developer of the evaluated frontier models

How it unfolded 3 developments, newest first · click a bar or a number to jump articlesposts

Peak 3 pieces in 3h at Sep 16, 12 PM; 11 pieces over 14 days (1 article · 3 posts · 7 comments) Sep 14, 3 AM — 1 piece · 1 post — Hacker News 1Sep 14, 6 AM — quietSep 14, 9 AM — quietSep 14, 12 PM — quietSep 14, 3 PM — quietSep 14, 6 PM — quietSep 14, 9 PM — quietSep 15, 12 AM — quietSep 15, 3 AM — quietSep 15, 6 AM — quietSep 15, 9 AM — quietSep 15, 12 PM — quietSep 15, 3 PM — quietSep 15, 6 PM — quietSep 15, 9 PM — quietSep 16, 12 AM — quietSep 16, 3 AM — quietSep 16, 6 AM — quietSep 16, 9 AM — quietSep 16, 12 PM — 3 pieces · 1 article · 1 post · 1 comment — Hacker News 2, Newswires 1Sep 16, 3 PM — 3 pieces · 3 comments — Hacker News 3Sep 16, 6 PM — 3 pieces · 3 comments — Hacker News 3Sep 16, 9 PM — 1 piece · 1 post — Hacker News 1Sep 17, 12 AM — quietSep 17, 3 AM — quietSep 17, 6 AM — quietSep 17, 9 AM — quietSep 17, 12 PM — quietSep 17, 3 PM — quietSep 17, 6 PM — quietSep 17, 9 PM — quietSep 18, 12 AM — quietSep 18, 3 AM — quietSep 18, 6 AM — quietSep 18, 9 AM — quietSep 18, 12 PM — quietSep 18, 3 PM — quietSep 18, 6 PM — quietSep 18, 9 PM — quietSep 19, 12 AM — quietSep 19, 3 AM — quietSep 19, 6 AM — quietSep 19, 9 AM — quietSep 19, 12 PM — quietSep 19, 3 PM — quietSep 19, 6 PM — quietSep 19, 9 PM — quietSep 20, 12 AM — quietSep 20, 3 AM — quietSep 20, 6 AM — quietSep 20, 9 AM — quietSep 20, 12 PM — quietSep 20, 3 PM — quietSep 20, 6 PM — quietSep 20, 9 PM — quietSep 21, 12 AM — quietSep 21, 3 AM — quietSep 21, 6 AM — quietSep 21, 9 AM — quietSep 21, 12 PM — quietSep 21, 3 PM — quietSep 21, 6 PM — quietSep 21, 9 PM — quietSep 22, 12 AM — quietSep 22, 3 AM — quietSep 22, 6 AM — quietSep 22, 9 AM — quietSep 22, 12 PM — quietSep 22, 3 PM — quietSep 22, 6 PM — quietSep 22, 9 PM — quietSep 23, 12 AM — quietSep 23, 3 AM — quietSep 23, 6 AM — quietSep 23, 9 AM — quietSep 23, 12 PM — quietSep 23, 3 PM — quietSep 23, 6 PM — quietSep 23, 9 PM — quietSep 24, 12 AM — quietSep 24, 3 AM — quietSep 24, 6 AM — quietSep 24, 9 AM — quietSep 24, 12 PM — quietSep 24, 3 PM — quietSep 24, 6 PM — quietSep 24, 9 PM — quietSep 25, 12 AM — quietSep 25, 3 AM — quietSep 25, 6 AM — quietSep 25, 9 AM — quietSep 25, 12 PM — quietSep 25, 3 PM — quietSep 25, 6 PM — quietSep 25, 9 PM — quietYesterday, 12 AM — quietYesterday, 3 AM — quietYesterday, 6 AM — quietYesterday, 9 AM — quietYesterday, 12 PM — quietYesterday, 3 PM — quietYesterday, 6 PM — quietYesterday, 9 PM — quietToday, 12 AM — quietToday, 3 AM — quietToday, 6 AM — quietToday, 9 AM — quietToday, 12 PM — quietToday, 3 PM — quietToday, 6 PM — quiet 123
Sep 15Sep 16Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25now · 8:27 PM ET
  1. 3

    Co-author jsous publishes 'Is Physics Dead' explainer

    One of the paper's authors publishes a blog post walking through the study's motivation and findings, arguing that despite low raw scores on CritPt (32% for GPT-5.6 Sol and GPT-6 Astra) and Humanity's Last Exam (47%), broken grading may be masking how close frontier models actually are to matching expert physics judgment.

    “Apparently, even a model good enough to be declared artificial general intelligence gets roughly half of graduate-level physics wrong.”
    — jsous
    • The last author also wrote up a quite-readable blogpost here that accompanies the article: https://jsous.github.io/blogs/is-physics-dead/One tidbit I found particularly interesting: "We tested a GPT-based agentic system, which previously succeeded in resolving several open mathematical conjectures, on open problems in theoretical physics. To our…

      quantumtwistHacker News10d agoview on Hacker News ↗
    2 more of the top 3 · 3 posts in this stretch
    • Starting to feel more and more like chinese room experimentThe models are confidently answering physics questions, treating it as a math problem, but they don't fundamentally "get it" and even recently failed simple "should i drive to car wash" testThe sample efficiency is just crazy lowStill surprising that even with this they managed to saturate…

      RomanKornevHacker News11d agoview on Hacker News ↗
    • Are models able to do math now, or do they still rely on “tools” to do the math?

      mch82Hacker News10d agoview on Hacker News ↗
    all of them →
  2. 2

    Paper gains traction on Hacker News front page

    The same arXiv paper is resubmitted and climbs to the Hacker News front page, drawing 85 points and 43 comments as the finding spreads beyond the initial small-scale post.

    • (Trained physicist here)From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced…

      amlutoHacker News11d agoview on Hacker News ↗
    2 more of the top 3 · 4 posts in this stretch
    • Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks"John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.When hand grading instead, they found out that the…

      qt31415926Hacker News11d agoview on Hacker News ↗
    • This is interesting and actually very important for robotics.I've been waiting for this, but all companies seem to not care much now.There is a way out of this by supplying right context (needs a bit of expertise in physics)1 more year and frontier will become crazy good at this as well.

      respectattentioHacker News11d agoview on Hacker News ↗
    all of them →
  3. 1 day quiet
  4. 1

    Researchers post physics benchmark re-grading study on arXiv

    A paper titled 'How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks' appears on arXiv, reporting that Yale physics faculty and graduate researchers manually audited how frontier models were scored on physics benchmarks.

    1. first by HN Frontpage, 11d ago

What people are saying 1 voices from 1 site · best of 7 · verbatim