Amazon Science publishes research on overfitting in ML research agents
1 Sep 14 · 13d ago · 1 article · 2 posts · 11 comments · 3 sources · development 1 of 1
Amazon Science releases an article examining why machine learning research avoids overfitting despite iteratively optimizing against benchmark datasets. The work uses LLM-based research agents to replicate human research community behavior in a controlled, resettable environment.
“Why is this being published as a blog post and not as a peer-reviewed submission? If it's going to be a blog post, why isn't there a corresponding scientific version for me to look at?”
jsrozner, Hacker News commenter · hn ↗Amazon Science Research organization
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by HN Frontpage, 13d ago
What people said 13 voices · best of 14 · verbatim
-
Compression in this modern day and age is so slop.Yes, I'm familiar with keystone results such as Solomonoff induction. It's a direct counterexample to compression - your intensional algorithm can completely outrun reality. I can literally specify a huge mega-algorithm that just searches over all possible Turing machines and evaluates them, and…
-
I think you should get less annoyed.> It’s not that the simplest is more likely to be correct, it’s that you should prefer it, because it’s simple.I don't know what Occam meant, but if you accept the formalism of PAC learning, it is more likely to be correct
-
That's not true. It's pretty clear that she meant "do something you knew they were going to say no to and now you are trying to get away with something."https://youtu.be/wHdHCoeUbU4?t=861s> So I want to tell something to all the young people here on many many occasions you'll find it is much easier to apologize than it is to get permission. You do…
-
Why is this being published as a blog post and not as a peer-reviewed submission? If it's going to be a blog post, why isn't there a corresponding scientific version for me to look at?Someone else already found it. I don't understand why the link isn't in the blog post. https://arxiv.org/abs/2606.11045Use of claude for writing it should be…
-
>It’s just like the Hopper quote.Not sure about Hopper, as I recall biographers of Lawrence of Arabia certainly made it seem like he was using the fog of war to do things he knew his superiors may object to.Regardless, even if its misinterpreted it still has a kernal of truth and separate utility than your version, that is: the people in the field…
-
“You should prefer it, because it’s simple” is just restating Occam’s razor, not giving any explanation. “The simplest is more likely to be accurate” is a much better interpretation than yours.I think the most accessible example of Occam’s razor is fitting a line to some points; you can always use a high enough order polynomial to fit the seen…
-
I always get annoyed when people misinterpret Occam’s razor. It’s not that the simplest is more likely to be correct, it’s that you should prefer it, because it’s simple.It’s just like the Hopper quote. She said it’s better to ask for forgiveness during the fog of war, doing something you thought was right, not to do something you knew they were…
-
If anything, the latest generation of AI models, Astra and Fable, are prime example of overfitting—whereas benchmarks suggest they’re AGI-tier, users (including myself) report the same old gaslighting, hallucination, context rot, cheating, incomprehensibility patterns as with prior models, sometimes even more pronounced.Fable and Opus 5, I…
-
Wherein Claude gives an honest assessment that it genuinely does not overfit. I also had Grok telling me that it isn't quantized.Do the submitters really not notice that this is AI slop? Do they like this? It is a complete pain to read.
-
simpler is not the right word either. it's the one that makes the least assumptions, not the simplest. The simplest would be "god did it" pretty much everytime.
-
And sadly, in academia, complexity (opposite of Occam's razor) is what gets you published.
-
Do you know what the scaling law actually is? Overfit everything as much as you can.
-
They tend not to overfit ... when there are way more data points than parameters.
All 1 developments of Amazon researchers study why ML research agents resist… →
Hacker NewsNewswiresLobstersMastodon