conv.

All stories
AIQuiet 25d · day 27

Researcher achieves 44% on ARC-AGI benchmark with $0.67 transformer

A small autoregressive transformer trained from scratch in 90 minutes outperforms larger language models on a reasoning benchmark at minimal cost.

What to know

  • A small transformer trained from scratch for 67 cents in 1.5 hours matched large language model performance on ARC-AGI-1, emphasizing sample efficiency over scale.
  • The approach uses test-time training on individual puzzles without access to evaluation labels, treating each puzzle's examples as training data during inference.
  • The work challenges the assumption that complex reasoning requires large models or expensive training, demonstrating architectural improvements and data efficiency as key factors.
  • Debate centers on whether using evaluation puzzle inputs (but not labels) for training constitutes a methodological flaw, with the researcher arguing it aligns with ARC's metalearning design.

The dispute Whether training on evaluation puzzle inputs (without labels) during test time constitutes data leakage or is a legitimate use of ARC's metalearning benchmark design. · positions read across 144 posts and comments

most voices

The result is technically impressive and validates sample efficiency as achievable without large models or high compute costs.

  • “First: this is really great technical writing, especially when you get into the rebuttals. Firm & clear without polemics -- props, and thanks for open-sourcing!”

    bbor · Hacker News
some voices

Using evaluation puzzle inputs at test time, even without labels, may constitute methodological flaw or 'cheating' that invalidates the benchmark comparison.

  • “Isn't this cheating? Or rather, are frontier agents only looking at one question at a time? If I understand correctly, you're looking at all the examples of the exam questions.”

    docheinestages · Hacker News

“I think sample efficiency is the most important problem in AI today and I want to solve it.”

porridgeraisin, AI researcher · Blog post

porridgeraisin / evilmathkid AI researcherLucas Beyer, Jeremy Howard, Rohan Anil Researchers (cited as discussing prior work)

The record 2 articles and posts · last 27 days

  1. 44% on ARC-AGI-1 in 67 cents press · HN Best · porridgeraisin · 26d ago · +1 outlet
  2. 44% on ARC-AGI-1 in 67 cents: https:// mvakde.github.io/blog/44-on-ar c-1/ Discussion: http://… post · Mastodon · newsyc250@toot.community · 26d ago

What people are saying 24 voices from 2 sites · best of 144 · verbatim

Still unanswered
  • If the model could only see one evaluation question at a time rather than all examples, would performance drop significantly?
  • Does spending 67 dollars instead of 67 cents lead to major performance improvements, or is there early saturation?