Researcher achieves 44% on ARC-AGI benchmark with $0.67 transformer
A small autoregressive transformer trained from scratch in 90 minutes outperforms larger language models on a reasoning benchmark at minimal cost.
What to know
- A small transformer trained from scratch for 67 cents in 1.5 hours matched large language model performance on ARC-AGI-1, emphasizing sample efficiency over scale.
- The approach uses test-time training on individual puzzles without access to evaluation labels, treating each puzzle's examples as training data during inference.
- The work challenges the assumption that complex reasoning requires large models or expensive training, demonstrating architectural improvements and data efficiency as key factors.
- Debate centers on whether using evaluation puzzle inputs (but not labels) for training constitutes a methodological flaw, with the researcher arguing it aligns with ARC's metalearning design.
The dispute Whether training on evaluation puzzle inputs (without labels) during test time constitutes data leakage or is a legitimate use of ARC's metalearning benchmark design. · positions read across 144 posts and comments
The result is technically impressive and validates sample efficiency as achievable without large models or high compute costs.
-
“First: this is really great technical writing, especially when you get into the rebuttals. Firm & clear without polemics -- props, and thanks for open-sourcing!”
bbor · Hacker News
Using evaluation puzzle inputs at test time, even without labels, may constitute methodological flaw or 'cheating' that invalidates the benchmark comparison.
-
“Isn't this cheating? Or rather, are frontier agents only looking at one question at a time? If I understand correctly, you're looking at all the examples of the exam questions.”
docheinestages · Hacker News
“I think sample efficiency is the most important problem in AI today and I want to solve it.”
porridgeraisin, AI researcher · Blog post
porridgeraisin / evilmathkid AI researcherLucas Beyer, Jeremy Howard, Rohan Anil Researchers (cited as discussing prior work)
The record 2 articles and posts · last 27 days
What people are saying 24 voices from 2 sites · best of 144 · verbatim
- If the model could only see one evaluation question at a time rather than all examples, would performance drop significantly?
- Does spending 67 dollars instead of 67 cents lead to major performance improvements, or is there early saturation?
- Sep 3
-
I didn't pull it out of my ass. It's basic logic that if you remove a supply restriction then the supply goes up.Let's attach fake numbers: Right now only 50 people can get residency each year, and 80 people starting college each year want to be doctors and would work hard enough. If you let everyone get residency, at first you'd make 80 doctors…
- Sep 1
-
In my unpopular opinion, it wasn't transformers or attention, but pretraining on language data that kicked off the LLM Boom. Alec radford in his little jupyter notebook trained a very small non-transformer to predict simply the next-character on Amazon reviews. He noticed emergence of a neuron which when toggled controlled the sentiment of the…
-
Hi,I loved reading your work and the blog post series and codebase.After doing the LLMs from scratch course because I had some free time on my hand, I reimplemented Karpathy's nanochat in Go with SIMD (as 1.27 finally supports it). Mostly to learn about how SIMD works, and how you can implement different kernels for your own model, how it works…
-
>Yes what you describe would be cheating. My approach is the opposite. What I did was "Carry your textbook to the exam and then learn from scratch during the exam">What I did was "You are born during the exam, given access to a training set (which is curated and allowed) and the questions then learn everything from scratch during the exam"This is…
-
I think you’re asking the right questions, sample inefficiency is horrible in modern LLMs. Despite this, I saw your analysis:> The biggest increases in scores were due toModern architecture (SwiGlu instead of GELU, RMSnorm not layernorm, etc.) More data diversity, better shuffling of data scaling up: 8 layers instead of 4This is commonly called…
-
It gets extremely blurry, because people commonly refer to any model that uses a component associated with the Transformer architecture as a Transformer (i.e. using some kind of QKV-esque attention mechanism). I think it's easier to think of it like this:A large language model is just what it says--a very large statistical model trained for…
-
Kind of hijacking, ... I'm glad you did! Of all the procrastination techniques I have mastered, engaging smart people on HN about artificial cognition is probably one of the more useful ;) Apologies in advance for the diatribe(s) -- I think about this stuff a lot. ...would you say that LLM's have solved the frame problem? To me, the frame problem…
-
> In the university I first dropped out of, students that surpassed me studied by getting and sharing copies of previous exams and solving those questionYes what you describe would be cheating. My approach is the opposite. What I did was "Carry your textbook to the exam and then learn from scratch during the exam"What you described is cheating…
-
Thanks!I think this is a great question. I have some thoughts on this but no hard evidence (neither does anyone else!)Your argument relies on the AGI system being the model arch + weights. I think that the weights are irrelevant. The training algorithm is what is AGI: You choose/find a training set that covers a task, and then some form of deep…
-
>Training on the eval puzzles is cheating / “training on test” No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on.I disagree, but i agree that training with answers is worse.In the university I first dropped out of, students that surpassed me studied by getting and sharing…
-
N
44% on ARC-AGI-1 in 67 cents: https:// mvakde.github.io/blog/44-on-ar c-1/ Discussion: http:// news.ycombinator.com/item?id=4 9519939
-
Kind of hijacking, would you say that LLM's have solved the frame problem?To me, the frame problem is: Can you function in an open vs closed world, and to me the answer is yes, LLM's can definitely function in an open world where the rules are fuzzy, changing, undefined, etc. At the very least, much better than all GOFAI approaches by far.The…
-
Yeah but rhabdo is something literally any e.g. body builder, power lifter, etc could tell you about. Actually if somebody knows what hypertrophy is, they probably know what rhabdo is. It's a pretty normal and big concern in any sort of high intensity weight training.I can't think of many ways that otherwise healthy and fit younger people can…
-
What you gather is correct, assuming by "the test" you mean the ARC benchmark in general. It was controversial because people are used to LLMs which are frozen at train time, where the eval problems are usually not trained on for various reasons like fragility (basically porridgeraisin's ans which is great)Here's another explanation. Take the…
-
First: this is really great technical writing, especially when you get into the rebuttals. Firm & clear without polemics -- props, and thanks for open-sourcing!That said; I don't have the time, energy, or anywhere near the expertise to challenge you on the DL specifics, but I feel compelled to add another voice to the chorus of doubters…
-
The point of ARC is essentially an "IQ Test" for AI systems. It is meant to cover abstract reasoning capabilities of generally-intelligent systems like LLMs. What the author did here was build a system that only solves ARC problems.The other tension is the fact that this score is on the public eval set. In machine learning, you typically have 3…
-
Basically, you have a bunch of Q,A pairs in the training dataset. Here, it was trained to next-word predict the question itself, as well as next-word predict the answer given the question as prompt. This is bog-standard, no one's complaining.In the test dataset's Q,A pairs, it was only trained to next-word predict the question itself, and it was…
-
I think(?) you’ve already probably done a good job of explaining this criticism for semi-informed people. But can you dumb it down even more for those of us who are almost entirely out-of-the-loop?> Training on the eval puzzles is cheating / “training on test”> No this is false. “Training on test” specifically means training on the labels of test…
-
To the parent posters credit, he's very honest that he developed a personal complex against doctors when he was a poor student. He just sees it as a way to lash out at American doctors, the irony being that foreign medical graduates leaving their families and communities to practice in America largely do so because they are exceptionally money…
-
I studied biochem in undergrad and my classes were full of premed students.I loved the subject and nerded out about the course material - I spent my time designing my own experiments around gene cloning that took several semesters to run. They were sharing last year's tests with their frat buddies and laughing at us nerds.I've never looked at…
-
How does it perform on ARC-AGI-3?There was this a few weeks ago:"Schema Harness Achieves ~99% on Arc‑AGI‑3 Public" https://news.ycombinator.com/item?id=48938163>> Schema, the harness we introduce today, reaches 99% on the ARC_AGI_3 Public set using Claude Opus 4.8 and Fable 5, and 95.35% using GPT‑5.6 SolWhat does that do with 5.6 Luna instead of…
-
Hi! Author here. Surprised to see this on HN now. Happy to answer any questions!Some context about this:- This is NOT an LLM. its a small ar transformer trained from scratch. One of the points was that extremely complex problems can be tackled without LLMs- Till the v1 of this result, this benchmark was only scaled by LLMs or their finetunes (ofc…
-
There are plenty of applications where a machine learning system needs to optimize for a very limited data set that is still intractable by linear logic systems of reasonable scale and complexity. It’s interesting, because he is using the legos of LLMs to build highly specialized machine learning systems, which is a very pragmatic approach…
-
> Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoningAgreed and that's for any benchmark. Private tests are better but you still have to trust the provider to not log and use them…