New papers push AI recursive self-improvement beyond math and code benchmarks
A cluster of research papers on AI systems that improve their own training pipeline surfaced on Hacker News within a single day.
What to know
- Multiple papers on AI 'recursive self-improvement' (RSI) reached Hacker News within roughly ten hours, suggesting a concentrated wave of research interest rather than one isolated finding.
- The most detailed proposal, MetaRSI-v1, combines three types of self-improvement operators (data, scaffold, model weights) but is only validated on standard coding and closed-form science benchmarks, not open-ended real-world tasks.
- The papers argue existing RSI work has been confined to machine-checkable domains like math and code, and call for extending it to broader scientific and engineering work where correctness is harder to verify.
- Evidence available is limited to paper abstracts and bare HN submission listings, with no substantive comment discussion captured.
Zihan Tan Co-lead author, MetaRSI-v1 paperLeixin Sun Co-lead author, MetaRSI-v1 paperGuancheng Wan Author, MetaRSI-v1 paper
How it unfolded 5 developments, newest first · click a bar or a number to jump posts
-
5
Curated 'awesome-rsi' GitHub list posted to HN
A GitHub repository collecting recursive self-improvement resources, 'awesome-rsi,' is shared on Hacker News shortly after the paper resubmission.
-
Perhaps I'm not understanding it correctly, but here's my take on what the paper is doing.Imagine you have a problem you want to solve (let's say, identify an OCR'd handwritten character, e.g. the MNIST Dataset). You tell 3 agents "Hey, each of you take a stab at getting really good at recognizing characters from this dataset. You can take 10…
2 more of the top 3 · 13 posts in this stretch
-
FYI; the paper is clearly a reference to Danijar Hafner's 'Dreamer' line of work, which was published in 2019, and which Danijar has continued to iterate on. https://arxiv.org/abs/1912.01603The TalkRL podcasts on this line of work are reasonable accessible and quite interesting.
-
Intuitively I wouldn't readjust how many steps they each do, but instead add another run afterwards, that get the same amount of steps as the previous, but now also with a concise description of what the previous attempts did and what they achieved, and ask it to improve. The amount of compute you have available, would dictate how many full…
-
-
4
'The Last AI Built by Humans' paper resurfaces on HN
The same 'Last AI Built by Humans' paper is resubmitted to Hacker News by a different user hours later, indicating renewed attention to the topic.
-
3
Researchers detail MetaRSI-v1 framework for AI self-improvement
A paper describing MetaRSI-v1 is posted, proposing a scheduled composition of three operators (Data-RSI, Harness-RSI, Model-RSI) that let a model improve data use, its own scaffold, and its parameters without external supervision; it is validated only on standard coding and closed-form science benchmarks.
“Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain.”
— MetaRSI-v1 paper authors -
2
'Meta^N' recursive self-improvement paper surfaces on HN
A separate arXiv paper, 'Meta$^N$: Recursive Self-Improvement Through Emergent Depth,' is posted to Hacker News the same evening.
-
1 outlet first by HN Frontpage, 11d ago · read ↗
-
first by arXiv cs.AI, 13d ago
-
-
1
Paper 'The Last AI Built by Humans' posted to arXiv
A paper titled 'The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement' is submitted to Hacker News, linking to its arXiv abstract page.
-
first by arXiv cs.AI, 12d ago
-
What people are saying 10 voices from 1 site · best of 13 · verbatim
- Sep 17
-
RSI has to be a harness because recursion is a loop, compute can be amortised by distilling behaviour into weights but the breakthrough would come from a harness. And agree, what Dream-RSI describes is a lot closer to online policy optimisation but pre-existence of a primitive doesn't have to detract from novel application.
- Sep 16
-
Is this a good place to ask why no one seems to be worried that recursive self improvement might be dangerous? To me that seems like a really bad idea but I’m interested to hear the pro-RSI side of things.
-
Your understanding is basically correct. "how applicable the search controller is when applied to new problems". We need meta-agent thinking pattern. Self-evolving agent has been very popular and we want to use agent to design a perfect agent. This is the problem that the "search" controller employed in this paper aims to solve.
-
Great explanation.Do you think this could be extrapolated to areas with no objectively verifiable results / outcomes?(Outside of math & science)
-
There is no way this could be reasonably framed as RSI.This iterative, online optimization of an exploration policy is not recursively intelligent in any way. It simply reallocates the available computational resources to more promising (hopefully) parts of the search space as system conditions change over time.
-
I don’t really understand this. Some real examples would help. Making a “simulator” out of a bunch of historic states does not tell me enough.
-
Isn't that a challenge with RL anyway that for a lot of problems its hard to even know accuracy continuously for each step
-
The replay simulator from history for off-policy eval is clever - avoids expensive rollouts. Curious how they prevent the policy from overfitting to already-discovered branches and going stale as the search space expands?
-
This is a solid and very interesting paper! The authors were kind enough to publish the complete prompt for it (appendix B.1 on page 18), so anyone can try their approach with any LLM and see the results.
-
Unless I'm misunderstanding, calling this RSI seems misleading?This looks like an optimization of current training methods, and a good one, but not "RSI" in the sense of a system that can perpetually improve itself forever.