Amazon researchers study why ML research agents resist overfitting
A new study uses LLM-based agents to test why benchmark-driven research produces real progress despite iterative optimization.
What to know
- ML research iteratively optimizes against reused benchmarks, yet produces real progress that transfers to new datasets—contradicting textbook predictions of rampant overfitting.
- Amazon researchers used LLM-based research agents that can be reset and controlled to test why this occurs, isolating variables that cannot be studied in human research communities.
- Comments note the paper was not linked in the blog post and question the suitability of blog format for scientific disclosure.
The dispute Whether the blog post and the underlying research represent genuine scientific insight or are examples of AI-generated overconfidence about models' capabilities. · positions read across 14 posts and comments
The blog post lacks rigor and transparency; it should link to peer-reviewed work and disclose if AI wrote it.
-
“Why is this being published as a blog post and not as a peer-reviewed submission? If it's going to be a blog post, why isn't there a corresponding scientific version for me to look at?”
jsrozner · Hacker News ↗
Current state-of-the-art models like Astra and Fable actually do overfit; they score high on benchmarks but fail in real use with hallucination and context degradation.
-
“If anything, the latest generation of AI models, Astra and Fable, are prime example of overfitting—whereas benchmarks suggest they're AGI-tier, users (including myself) report the same old gaslighting, hallucination, context rot, cheating…”
ubutler · Hacker News ↗
The paper's premise about overfitting resistance may hold only in data-rich regimes where parameters are far fewer than data points.
-
“They tend not to overfit ... when there are way more data points than parameters.”
nyeah · Hacker News ↗
“Studies that build entirely fresh test sets for old, heavily reused benchmarks have found that improvements largely transfer: on the new data, models demonstrate the same gains they did on the old benchmark.”
Amazon Science · Amazon Science blog
Amazon Science Research organization
How it unfolded 1 development · click the chart to see its coverage articlespostscomments
-
1
“Why is this being published as a blog post and not as a peer-reviewed submission? If it's going to be a blog post, why isn't there a corresponding scientific version for me to look at?”
— jsrozner, Hacker News commenter · source -
first by HN Frontpage, 13d ago
-
Compression in this modern day and age is so slop.Yes, I'm familiar with keystone results such as Solomonoff induction. It's a direct counterexample to compression - your intensional algorithm can completely outrun reality. I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them, and…
2 more of the top 3 · 14 posts in this stretch
-
I think you should get less annoyed.> It’s not that the simplest is more likely to be correct, it’s that you should prefer it, because it’s simple.I don't know what Occam meant, but if you accept the formalism of PAC learning, it is more likely to be correct
-
That's not true. It's pretty clear that she meant "do something you knew they were going to say no to and now you are trying to get away with something."https://youtu.be/wHdHCoeUbU4?t=861s> So I want to tell something to all the young people here on many many occasions you'll find it is much easier to apologize than it is to get permission. You do…
-
What people are saying 9 voices from 1 site · best of 14 · verbatim
- Sep 15
-
simpler is not the right word either. it's the one that makes the least assumptions, not the simplest. The simplest would be "god did it" pretty much everytime.
-
Do you know what the scaling law actually is? Overfit everything as much as you can.
- Sep 14
-
“You should prefer it, because it’s simple” is just restating Occam’s razor, not giving any explanation. “The simplest is more likely to be accurate” is a much better interpretation than yours.I think the most accessible example of Occam’s razor is fitting a line to some points; you can always use a high enough order polynomial to fit the seen…
-
If anything, the latest generation of AI models, Astra and Fable, are prime example of overfitting—whereas benchmarks suggest they’re AGI-tier, users (including myself) report the same old gaslighting, hallucination, context rot, cheating, incomprehensibility patterns as with prior models, sometimes even more pronounced.Fable and Opus 5, I…
-
And sadly, in academia, complexity (opposite of Occam's razor) is what gets you published.
-
>It’s just like the Hopper quote.Not sure about Hopper, as I recall biographers of Lawrence of Arabia certainly made it seem like he was using the fog of war to do things he knew his superiors may object to.Regardless, even if its misinterpreted it still has a kernal of truth and separate utility than your version, that is: the people in the field…
-
Wherein Claude gives an honest assessment that it genuinely does not overfit. I also had Grok telling me that it isn't quantized.Do the submitters really not notice that this is AI slop? Do they like this? It is a complete pain to read.
-
They tend not to overfit ... when there are way more data points than parameters.
-
I always get annoyed when people misinterpret Occam’s razor. It’s not that the simplest is more likely to be correct, it’s that you should prefer it, because it’s simple.It’s just like the Hopper quote. She said it’s better to ask for forgiveness during the fog of war, doing something you thought was right, not to do something you knew they were…