conv.

All stories
AIQuiet 13d · day 14

Amazon researchers study why ML research agents resist overfitting

A new study uses LLM-based agents to test why benchmark-driven research produces real progress despite iterative optimization.

What to know

  • ML research iteratively optimizes against reused benchmarks, yet produces real progress that transfers to new datasets—contradicting textbook predictions of rampant overfitting.
  • Amazon researchers used LLM-based research agents that can be reset and controlled to test why this occurs, isolating variables that cannot be studied in human research communities.
  • Comments note the paper was not linked in the blog post and question the suitability of blog format for scientific disclosure.

The dispute Whether the blog post and the underlying research represent genuine scientific insight or are examples of AI-generated overconfidence about models' capabilities. · positions read across 14 posts and comments

many voices

The blog post lacks rigor and transparency; it should link to peer-reviewed work and disclose if AI wrote it.

  • “Why is this being published as a blog post and not as a peer-reviewed submission? If it's going to be a blog post, why isn't there a corresponding scientific version for me to look at?”

    jsrozner · Hacker News ↗
some voices

Current state-of-the-art models like Astra and Fable actually do overfit; they score high on benchmarks but fail in real use with hallucination and context degradation.

  • “If anything, the latest generation of AI models, Astra and Fable, are prime example of overfitting—whereas benchmarks suggest they're AGI-tier, users (including myself) report the same old gaslighting, hallucination, context rot, cheating…”

    ubutler · Hacker News ↗
some voices

The paper's premise about overfitting resistance may hold only in data-rich regimes where parameters are far fewer than data points.

  • “They tend not to overfit ... when there are way more data points than parameters.”

    nyeah · Hacker News ↗

“Studies that build entirely fresh test sets for old, heavily reused benchmarks have found that improvements largely transfer: on the new data, models demonstrate the same gains they did on the old benchmark.”

Amazon Science · Amazon Science blog

Amazon Science Research organization

How it unfolded 1 development · click the chart to see its coverage articlespostscomments

Peak 11 pieces in 3h at Sep 14, 12 PM; 18 pieces over 14 days (1 article · 3 posts · 14 comments) Sep 14, 12 PM — 11 pieces · 1 article · 2 posts · 8 comments — Hacker News 9, Newswires 1, Lobsters 1Sep 14, 3 PM — 2 pieces · 2 comments — Hacker News 2Sep 14, 6 PM — 2 pieces · 1 post · 1 comment — Hacker News 1, Mastodon 1Sep 14, 9 PM — 1 piece · 1 comment — Hacker News 1Sep 15, 12 AM — 1 piece · 1 comment — Hacker News 1Sep 15, 3 AM — quietSep 15, 6 AM — 1 piece · 1 comment — Hacker News 1Sep 15, 9 AM — quietSep 15, 12 PM — quietSep 15, 3 PM — quietSep 15, 6 PM — quietSep 15, 9 PM — quietSep 16, 12 AM — quietSep 16, 3 AM — quietSep 16, 6 AM — quietSep 16, 9 AM — quietSep 16, 12 PM — quietSep 16, 3 PM — quietSep 16, 6 PM — quietSep 16, 9 PM — quietSep 17, 12 AM — quietSep 17, 3 AM — quietSep 17, 6 AM — quietSep 17, 9 AM — quietSep 17, 12 PM — quietSep 17, 3 PM — quietSep 17, 6 PM — quietSep 17, 9 PM — quietSep 18, 12 AM — quietSep 18, 3 AM — quietSep 18, 6 AM — quietSep 18, 9 AM — quietSep 18, 12 PM — quietSep 18, 3 PM — quietSep 18, 6 PM — quietSep 18, 9 PM — quietSep 19, 12 AM — quietSep 19, 3 AM — quietSep 19, 6 AM — quietSep 19, 9 AM — quietSep 19, 12 PM — quietSep 19, 3 PM — quietSep 19, 6 PM — quietSep 19, 9 PM — quietSep 20, 12 AM — quietSep 20, 3 AM — quietSep 20, 6 AM — quietSep 20, 9 AM — quietSep 20, 12 PM — quietSep 20, 3 PM — quietSep 20, 6 PM — quietSep 20, 9 PM — quietSep 21, 12 AM — quietSep 21, 3 AM — quietSep 21, 6 AM — quietSep 21, 9 AM — quietSep 21, 12 PM — quietSep 21, 3 PM — quietSep 21, 6 PM — quietSep 21, 9 PM — quietSep 22, 12 AM — quietSep 22, 3 AM — quietSep 22, 6 AM — quietSep 22, 9 AM — quietSep 22, 12 PM — quietSep 22, 3 PM — quietSep 22, 6 PM — quietSep 22, 9 PM — quietSep 23, 12 AM — quietSep 23, 3 AM — quietSep 23, 6 AM — quietSep 23, 9 AM — quietSep 23, 12 PM — quietSep 23, 3 PM — quietSep 23, 6 PM — quietSep 23, 9 PM — quietSep 24, 12 AM — quietSep 24, 3 AM — quietSep 24, 6 AM — quietSep 24, 9 AM — quietSep 24, 12 PM — quietSep 24, 3 PM — quietSep 24, 6 PM — quietSep 24, 9 PM — quietSep 25, 12 AM — quietSep 25, 3 AM — quietSep 25, 6 AM — quietSep 25, 9 AM — quietSep 25, 12 PM — quietSep 25, 3 PM — quietSep 25, 6 PM — quietSep 25, 9 PM — quietSep 26, 12 AM — quietSep 26, 3 AM — quietSep 26, 6 AM — quietSep 26, 9 AM — quietSep 26, 12 PM — quietSep 26, 3 PM — quietSep 26, 6 PM — quietSep 26, 9 PM — quietYesterday, 12 AM — quietYesterday, 3 AM — quietYesterday, 6 AM — quietYesterday, 9 AM — quietYesterday, 12 PM — quietYesterday, 3 PM — quietYesterday, 6 PM — quietYesterday, 9 PM — quietToday, 12 AM — quiet 1
Sep 15Sep 16Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25Sep 26now · 1:13 AM ET
  1. 1
    “Why is this being published as a blog post and not as a peer-reviewed submission? If it's going to be a blog post, why isn't there a corresponding scientific version for me to look at?”
    — jsrozner, Hacker News commenter · source
    1. first by HN Frontpage, 13d ago

    • Compression in this modern day and age is so slop.Yes, I'm familiar with keystone results such as Solomonoff induction. It's a direct counterexample to compression - your intensional algorithm can completely outrun reality. I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them, and…

      sigbottleHacker News13d agoview on Hacker News ↗
    2 more of the top 3 · 14 posts in this stretch
    • I think you should get less annoyed.> It’s not that the simplest is more likely to be correct, it’s that you should prefer it, because it’s simple.I don't know what Occam meant, but if you accept the formalism of PAC learning, it is more likely to be correct

      sreanHacker News13d agoview on Hacker News ↗
    • That's not true. It's pretty clear that she meant "do something you knew they were going to say no to and now you are trying to get away with something."https://youtu.be/wHdHCoeUbU4?t=861s> So I want to tell something to all the young people here on many many occasions you'll find it is much easier to apologize than it is to get permission. You do…

      gowldHacker News13d agoview on Hacker News ↗
    all of them →

What people are saying 9 voices from 1 site · best of 14 · verbatim