Entelligence benchmark: cheap GPT-5.6 Luna catches 75% of bugs Astra finds, for 3.6% of cost
A new code-review benchmark pits a $1.20-per-million-token model against OpenAI's premium GPT-6 Astra on 50 real pull requests.
What to know
- Luna costs $0.20/$1.20 per million input/output tokens vs Astra's $10/$50, a 28x per-review cost gap.
- Luna found 69 verified bugs to Astra's 92 (75%) at 3.6% of the total cost across 50 benchmark pull requests.
- Luna's accuracy drops sharply on security-sensitive code: only 50% of its Keycloak (identity/auth) findings verified, versus 93% for Astra.
- Entelligence recommends Luna for routine correctness bugs but not for unsupervised review of authentication or permission logic.
Entelligence.ai AI code-review benchmarking companyGPT-5.6 Luna Low-cost coding model under testGPT-6 Astra Premium coding model under testGPT-5.6 Sol Secondary verification judge / prior benchmark comparator
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Post reaches Hacker News front page
The benchmark writeup is submitted to Hacker News, drawing a score of 55 and 64 comments, indicating broader developer interest in the cost/accuracy tradeoff of AI code review models.
“Our read: Luna is good enough for everyday correctness bugs at that price, and we wouldn't let it review authentication or permission code on its own.”
— Entelligence.ai, benchmark author · source -
I was using Copilot Code Review pretty religiously for a while, as I get access for free (the $10 plan) due to my Open Source work, but it recently introduced a monster of a misfeature that caused a massive increase in complexity over time, while I wasn't paying close enough attention to it. Every subsequent model saw that change and the…
2 more of the top 3 · 24 posts in this stretch
-
I found Luna and even 5.4-mini to be quite good at code review provided a few things:1. Run it in multiple cycles, only on the diff, and only emit a few findings at a time.2. Give it a memory so each cycle, it knows the previous finding to check if it's been fixed.3. Give it access to canonical docs that encode your human reviewer heuristics. I…
-
In our new world of non-deterministic output (that's why we love LLMs! they say such helpful/agreeable/sometimes wrong stuff!), I think CI won't be sufficient. CI is in the realm of Quality Control; when I build the thing, is it to spec and does it do what I need it to do?But when the model can shift underneath you, I think it will put pressure on…
-
-
background
Entelligence publishes full Luna vs Astra benchmark writeup — The full blog post details methodology: same prompt on same diffs, dual-judge verification by Astra and GPT-5.6 Sol, per-codebase and per-bug-class breakdowns across Cal.com, Sentry, Discourse, Keycloak and Grafana repositories.
-
1
Entelligence posts Luna vs Astra code-review benchmark on Reddit
Entelligence.ai's Reddit account summarizes a new benchmark comparing GPT-5.6 Luna and GPT-6 Astra on 50 pull requests, reporting Luna caught 75% of Astra's confirmed bugs for 3.6% of the cost, and previews an upcoming Astra vs Fable 5.1 comparison.
“Astra found 92 confirmed bugs vs 69 for Luna, while Luna caught 75% of the bugs at just 3.6% of the cost…”
— entelligenceai17 -
first by HN Frontpage, 13d ago
-
What people are saying 21 voices from 1 site · best of 24 · verbatim
- Sep 15
-
Updated link: https://entelligence.ai/blogs/gpt-6-astra-cost-1.6x-more-per...The current link is a 404, seems like they didn't redirect it properly.
-
Luna is great for grunt work, and has some intelligence, but the context window makes it terrible for any decent-sized review.
-
Knowing what model to use depending on the goal is the best practice using agentic AI.
-
> The pull requests are public, and the prompts, raw model outputs, judge verdicts, bug-class labels, repeat runs and scoring scripts are committed with this article.Okay, but where are they? This quote says that PRs are public. Does that mean that everything else is private and we can’t actually reproduce those results?This, and the fact that the…
- Sep 14
-
We’ve added AI to our auto review process. It does expose when the author is not confident in their solution to push back. Which is interesting. But we do try to target specifically at our patterns and safety.
-
Nah, we have AI code review at Google and it is shockingly good at catching bugs no one would have noticed. I absolutely depend on it now.
-
For the mj adversarial review we use a configurable coordinator (I recommend Fable) with six subagents (I recommend Luna) for specific specialties like code complexity. Narrowing down what a smaller model like Luna needs to look for / care about helps them do better work; then the larger model synthesizes and fills in any gaps.
-
This article is missing an incredibly important detail: what is the harness doing?I get remarkably good results using any recent OpenAI model using the codex-rs harness pointing at a built checkout of the PR. The models use the available tools (i.e. the shell) to understand the repo. I get some false positives and some false negatives, but I don’t…
-
Any model is good enough. Even tiny local models can provide some value and their false positive rate is still relatively low and warrants a proper reply.
-
Our agents automatically review our PR's - the authors agents automatically see the feedback and make fixes, and automatically merge when everything is green.A well authored CI review process is significantly better than any human could do. We have the AI review not only the changes but clone and investigate all related repositories that integrate…
-
I'm glad I work in places where there's no such silly pointless rules like how many people need to review a PR.The PR author asks for feedback if it needs feedback, otherwise it merges it, period.I don't know why and when the world got convinced that all this bureaucracy is a "best practice", when it's just a practice, that can be good, or a waste…
-
I select my review model based on change complexity. If I'm confident that the change is localized (given that I always maintain an up to date and complete mental model of the application), I'll use Luna/Terra. For deeper changes, I use Sol/Astra. Sol is so good for code reviews that I'll only engage Astra in the very riskiest changes.
-
They state Luna is good enough, but its accuracy of findings is 74% whereas Astra is 96%. Dealing with false positives is expensive.I am finding AI doing its own reviews as part of the process to be the key to productivity. I do subagent (fresh context reviews) at multiple stages with well-specified review criteria. It is really expensive to do…
-
False positives have a real cost, especially if AI is reading a review. Consider if you have GPT-6 Astra looking at a review and finding a bunch of false positives it burns tokens to figure out.
-
We had a two human PR requirement until recently we dropped it. It was slowing us down too much now the human developer creating the future is obviously writing it all with AI so they need to check it then depending on the feature and it’s use it requires a PR but it’s not universal and we’ve stepped up our automated test Tan X what it used to be…
-
In my experience with the Claude Github integration, I found it to be pretty helpful. It’s had a pretty good success rate of catching bugs before they get to master, and for simple ones I can ask it to fix itself.> you should have 2+ developers looking at most PRsIt’d be nice, but usually not the case in my experience. More eyes is better. AI…
-
> You wouldn't ask an agent to review a PR then just copy/paste the output int PR would you?Hasn't everyone already got agents directly adding themselves to PRs and leaving comments (occasionally useful)?
-
AI should be used for code review but not in CI.You should already have 2+ developers looking at most PRs. And these developers should absolutely use AI. The PR author should use AI.But what you should not do is pipe the AI output directly into the PR and tell the PR author to deal with it. That's adding noise to the PR review process. Everything…
-
I only use chinese models for code reviews because you can actually tell them to take an adversarial stance and actively look for security issues without risking refusals. GLM-5.3 has been great for this, although it can be slow on larger PRs.
-
IMHO, Codex with Astra/Sol and Claude with Fable/Opus are all any professional programmer should be using in Sep 2026, if they can afford it.These models are still terrible compared to what we'd actually wish for, but they're the best available.If you can get away with using the $200/mo subscriptions, it's really not even a money thing for most…
-
$0.10 extra per pr review is nothing. What software company is willing to accept worse reviews and less bugs found to save 10 cents?