Paper gains traction on Hacker News front page
2 Sep 16 3:19 PM · 11d ago · 1 article · 1 post · 2 sources · development 2 of 3
The same arXiv paper is resubmitted and climbs to the Hacker News front page, drawing 85 points and 43 comments as the finding spreads beyond the initial small-scale post.
jsous Paper co-author and blog post authorArman Cohan Co-author of the benchmark re-grading studyYale physics faculty and graduate researchers Conducted expert audits of benchmark gradingOpenAI Developer of the evaluated frontier models
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by HN Frontpage, 11d ago
What people said 4 voices · verbatim
-
(Trained physicist here)From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced…
-
Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks"John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.When hand grading instead, they found out that the…
-
This is interesting and actually very important for robotics.I've been waiting for this, but all companies seem to not care much now.There is a way out of this by supplying right context (needs a bit of expertise in physics)1 more year and frontier will become crazy good at this as well.
-
I would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.
All 3 developments of Yale-led audit finds physics AI benchmarks are broken, not… →
Hacker NewsNewswires