conv.

All stories
AIQuiet 13d · day 14

Entelligence benchmark: cheap GPT-5.6 Luna catches 75% of bugs Astra finds, for 3.6% of cost

A new code-review benchmark pits a $1.20-per-million-token model against OpenAI's premium GPT-6 Astra on 50 real pull requests.

What to know

  • Luna costs $0.20/$1.20 per million input/output tokens vs Astra's $10/$50, a 28x per-review cost gap.
  • Luna found 69 verified bugs to Astra's 92 (75%) at 3.6% of the total cost across 50 benchmark pull requests.
  • Luna's accuracy drops sharply on security-sensitive code: only 50% of its Keycloak (identity/auth) findings verified, versus 93% for Astra.
  • Entelligence recommends Luna for routine correctness bugs but not for unsupervised review of authentication or permission logic.

Entelligence.ai AI code-review benchmarking companyGPT-5.6 Luna Low-cost coding model under testGPT-6 Astra Premium coding model under testGPT-5.6 Sol Secondary verification judge / prior benchmark comparator

Entelligence benchmark: cheap GPT-5.6 Luna catches 75% of bugs Astra finds, for 3.6% of cost
entelligence.ai

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 11 pieces in 3h at Sep 14, 1 PM; 27 pieces over 14 days (1 article · 2 posts · 24 comments) Sep 14, 10 AM — 1 piece · 1 post — Reddit 1Sep 14, 1 PM — 11 pieces · 1 article · 1 post · 9 comments — Hacker News 10, Newswires 1Sep 14, 4 PM — 9 pieces · 9 comments — Hacker News 9Sep 14, 7 PM — 2 pieces · 2 comments — Hacker News 2Sep 14, 10 PM — 1 piece · 1 comment — Hacker News 1Sep 15, 1 AM — quietSep 15, 4 AM — quietSep 15, 7 AM — 1 piece · 1 comment — Hacker News 1Sep 15, 10 AM — 2 pieces · 2 comments — Hacker News 2Sep 15, 1 PM — quietSep 15, 4 PM — quietSep 15, 7 PM — quietSep 15, 10 PM — quietSep 16, 1 AM — quietSep 16, 4 AM — quietSep 16, 7 AM — quietSep 16, 10 AM — quietSep 16, 1 PM — quietSep 16, 4 PM — quietSep 16, 7 PM — quietSep 16, 10 PM — quietSep 17, 1 AM — quietSep 17, 4 AM — quietSep 17, 7 AM — quietSep 17, 10 AM — quietSep 17, 1 PM — quietSep 17, 4 PM — quietSep 17, 7 PM — quietSep 17, 10 PM — quietSep 18, 1 AM — quietSep 18, 4 AM — quietSep 18, 7 AM — quietSep 18, 10 AM — quietSep 18, 1 PM — quietSep 18, 4 PM — quietSep 18, 7 PM — quietSep 18, 10 PM — quietSep 19, 1 AM — quietSep 19, 4 AM — quietSep 19, 7 AM — quietSep 19, 10 AM — quietSep 19, 1 PM — quietSep 19, 4 PM — quietSep 19, 7 PM — quietSep 19, 10 PM — quietSep 20, 1 AM — quietSep 20, 4 AM — quietSep 20, 7 AM — quietSep 20, 10 AM — quietSep 20, 1 PM — quietSep 20, 4 PM — quietSep 20, 7 PM — quietSep 20, 10 PM — quietSep 21, 1 AM — quietSep 21, 4 AM — quietSep 21, 7 AM — quietSep 21, 10 AM — quietSep 21, 1 PM — quietSep 21, 4 PM — quietSep 21, 7 PM — quietSep 21, 10 PM — quietSep 22, 1 AM — quietSep 22, 4 AM — quietSep 22, 7 AM — quietSep 22, 10 AM — quietSep 22, 1 PM — quietSep 22, 4 PM — quietSep 22, 7 PM — quietSep 22, 10 PM — quietSep 23, 1 AM — quietSep 23, 4 AM — quietSep 23, 7 AM — quietSep 23, 10 AM — quietSep 23, 1 PM — quietSep 23, 4 PM — quietSep 23, 7 PM — quietSep 23, 10 PM — quietSep 24, 1 AM — quietSep 24, 4 AM — quietSep 24, 7 AM — quietSep 24, 10 AM — quietSep 24, 1 PM — quietSep 24, 4 PM — quietSep 24, 7 PM — quietSep 24, 10 PM — quietSep 25, 1 AM — quietSep 25, 4 AM — quietSep 25, 7 AM — quietSep 25, 10 AM — quietSep 25, 1 PM — quietSep 25, 4 PM — quietSep 25, 7 PM — quietSep 25, 10 PM — quietSep 26, 1 AM — quietSep 26, 4 AM — quietSep 26, 7 AM — quietSep 26, 10 AM — quietSep 26, 1 PM — quietSep 26, 4 PM — quietSep 26, 7 PM — quietSep 26, 10 PM — quietYesterday, 1 AM — quietYesterday, 4 AM — quietYesterday, 7 AM — quietYesterday, 10 AM — quietYesterday, 1 PM — quietYesterday, 4 PM — quietYesterday, 7 PM — quietYesterday, 10 PM — quiet 1–2
Sep 15Sep 16Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25Sep 26now · 1:54 AM ET
  1. 2

    Post reaches Hacker News front page

    The benchmark writeup is submitted to Hacker News, drawing a score of 55 and 64 comments, indicating broader developer interest in the cost/accuracy tradeoff of AI code review models.

    “Our read: Luna is good enough for everyday correctness bugs at that price, and we wouldn't let it review authentication or permission code on its own.”
    — Entelligence.ai, benchmark author · source
    • I was using Copilot Code Review pretty religiously for a while, as I get access for free (the $10 plan) due to my Open Source work, but it recently introduced a monster of a misfeature that caused a massive increase in complexity over time, while I wasn't paying close enough attention to it. Every subsequent model saw that change and the…

      SwellJoeHacker News13d agoview on Hacker News ↗
    2 more of the top 3 · 24 posts in this stretch
    • I found Luna and even 5.4-mini to be quite good at code review provided a few things:1. Run it in multiple cycles, only on the diff, and only emit a few findings at a time.2. Give it a memory so each cycle, it knows the previous finding to check if it's been fixed.3. Give it access to canonical docs that encode your human reviewer heuristics. I…

      CharlieDigitalHacker News13d agoview on Hacker News ↗
    • In our new world of non-deterministic output (that's why we love LLMs! they say such helpful/agreeable/sometimes wrong stuff!), I think CI won't be sufficient. CI is in the realm of Quality Control; when I build the thing, is it to spec and does it do what I need it to do?But when the model can shift underneath you, I think it will put pressure on…

      gavinbostonHacker News13d agoview on Hacker News ↗
    all of them →
  2. background

    Entelligence publishes full Luna vs Astra benchmark writeup — The full blog post details methodology: same prompt on same diffs, dual-judge verification by Astra and GPT-5.6 Sol, per-codebase and per-bug-class breakdowns across Cal.com, Sentry, Discourse, Keycloak and Grafana repositories.

  3. 1

    Entelligence posts Luna vs Astra code-review benchmark on Reddit

    Entelligence.ai's Reddit account summarizes a new benchmark comparing GPT-5.6 Luna and GPT-6 Astra on 50 pull requests, reporting Luna caught 75% of Astra's confirmed bugs for 3.6% of the cost, and previews an upcoming Astra vs Fable 5.1 comparison.

    “Astra found 92 confirmed bugs vs 69 for Luna, while Luna caught 75% of the bugs at just 3.6% of the cost…”
    — entelligenceai17
    1. first by HN Frontpage, 13d ago

What people are saying 21 voices from 1 site · best of 24 · verbatim