Curated 'awesome-rsi' GitHub list posted to HN
5 Sep 14 4:17 AM · 13d ago · 1 post · 1 source · development 5 of 5
A GitHub repository collecting recursive self-improvement resources, 'awesome-rsi,' is shared on Hacker News shortly after the paper resubmission.
Zihan Tan Co-lead author, MetaRSI-v1 paperLeixin Sun Co-lead author, MetaRSI-v1 paperGuancheng Wan Author, MetaRSI-v1 paper
The whole story posts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 4 claims were published in its stretch
-
first by HN Frontpage, 11d ago
-
first by arXiv cs.AI, 12d ago
-
first by arXiv cs.AI, 13d ago
-
first by arXiv cs.AI, 13d ago
What people said 13 voices · verbatim
-
Perhaps I'm not understanding it correctly, but here's my take on what the paper is doing.Imagine you have a problem you want to solve (let's say, identify an OCR'd handwritten character, e.g. the MNIST Dataset). You tell 3 agents "Hey, each of you take a stab at getting really good at recognizing characters from this dataset. You can take 10…
-
FYI; the paper is clearly a reference to Danijar Hafner's 'Dreamer' line of work, which was published in 2019, and which Danijar has continued to iterate on. https://arxiv.org/abs/1912.01603The TalkRL podcasts on this line of work are reasonable accessible and quite interesting.
-
Intuitively I wouldn't readjust how many steps they each do, but instead add another run afterwards, that get the same amount of steps as the previous, but now also with a concise description of what the previous attempts did and what they achieved, and ask it to improve. The amount of compute you have available, would dictate how many full…
-
Your understanding is basically correct. "how applicable the search controller is when applied to new problems". We need meta-agent thinking pattern. Self-evolving agent has been very popular and we want to use agent to design a perfect agent. This is the problem that the "search" controller employed in this paper aims to solve.
-
RSI has to be a harness because recursion is a loop, compute can be amortised by distilling behaviour into weights but the breakthrough would come from a harness. And agree, what Dream-RSI describes is a lot closer to online policy optimisation but pre-existence of a primitive doesn't have to detract from novel application.
-
There is no way this could be reasonably framed as RSI.This iterative, online optimization of an exploration policy is not recursively intelligent in any way. It simply reallocates the available computational resources to more promising (hopefully) parts of the search space as system conditions change over time.
-
Unless I'm misunderstanding, calling this RSI seems misleading?This looks like an optimization of current training methods, and a good one, but not "RSI" in the sense of a system that can perpetually improve itself forever.
-
The replay simulator from history for off-policy eval is clever - avoids expensive rollouts. Curious how they prevent the policy from overfitting to already-discovered branches and going stale as the search space expands?
-
Is this a good place to ask why no one seems to be worried that recursive self improvement might be dangerous? To me that seems like a really bad idea but I’m interested to hear the pro-RSI side of things.
-
This is a solid and very interesting paper! The authors were kind enough to publish the complete prompt for it (appendix B.1 on page 18), so anyone can try their approach with any LLM and see the results.
-
Great explanation.Do you think this could be extrapolated to areas with no objectively verifiable results / outcomes?(Outside of math & science)
-
I don’t really understand this. Some real examples would help. Making a “simulator” out of a bunch of historic states does not tell me enough.
-
Isn't that a challenge with RL anyway that for a lot of problems its hard to even know accuracy continuously for each step
All 5 developments of New papers push AI recursive self-improvement beyond math… →
NewswiresHacker NewsMastodon