AIQuiet 12d · day 13
Researchers identify Matthew Effect bias in LLM reinforcement learning
A new paper shows RL training improves easy problems while hard ones stagnate, proposing Never Give Up as a solution.
What to know
- RL training of LLMs exhibits a Matthew Effect: gains concentrate on problems the model already solves well, while hard problems remain difficult.
- The bias appears across math, code, and agentic reasoning tasks, not just in specific domains.
- Researchers propose Never Give Up as a technique to address this cumulative advantage pattern in reinforcement learning.
Michael Noukhovitch Lead researcherHamish Ivison Co-authorNathan Lambert Co-authorAaron Courville Co-author
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
Sep 15Sep 16Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 11:07 PM ET
-
2
Blog post released explaining Matthew Effect and Never Give Up technique
Michael Noukhovitch published an interactive blog post accompanying the arxiv paper, breaking down the phenomenon with eval curves from Olmo 3.1 RL-Zero Math training on AIME 2025 questions.
“The majority of our improvements are coming from the easiest problems going from somewhat solved to mostly solved. The hardest problems are barely improving.”
— Michael Noukhovitch -
1
Noukhovitch et al. publish Matthew Effect paper on arXiv
A research paper titled "Learning to Solve Hard Problems in RL for LLMs by Never Giving Up" was posted to arXiv, introducing the concept of the Matthew Effect in RL for LLMs and proposing a solution method.
-
first by HN Frontpage, 12d ago · also arXiv cs.AI
-