GPT-6 Astra successfully drives a real car in benchmark test
A cloud-based language model completes a cone course in a Toyota Corolla, sparking debate over whether LLMs can meaningfully replace specialized autonomous driving systems.
What to know
- GPT-6 Astra successfully completed a real-world driving task—navigating a cone course in a Toyota Corolla—demonstrating frontier LLMs can achieve spatial and temporal reasoning out-of-the-box.
- Cloud latency is the disqualifying barrier to practical deployment; the benchmark operates at extremely low speeds and lacks the 20Hz update frequency that production autonomous systems require.
- The achievement reignites debate over whether generalist LLMs with strong vision capabilities could eventually replace specialized autonomous driving architectures, though near-term viability remains disputed.
- Benchmark creators frame this as a capability test, not a deployable system; commenters are divided on whether the result signals a genuine architectural shift or is merely a novelty application of large models.
The dispute Whether this benchmark represents a genuine architectural shift toward LLM-driven autonomy or is merely a novelty application of generalist models that falls short of practical deployment due to latency and speed constraints. · positions read across 21 posts and comments
Cloud LLMs could reshape autonomous driving by replacing specialized vision stacks with end-to-end generalist models.
-
“The bitter lesson is finally coming for the self-driving cars… it's maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.”
valine · Hacker News ↗
Cloud latency makes this benchmark impractical for real-world driving; specialized local systems remain superior.
-
“Will this work in the real world? Absolutely not. Three reasons: latency, latency, and latency. openpilot's driving model updates the target curvature and acceleration at 20Hz.”
jyoung8607 · Hacker News ↗
A hybrid architecture—local hardware for low-level control, cloud for high-level reasoning—could combine the strengths of both approaches.
-
“Eyes, control, and safety critical features on the hardware, higher level decision making to the cloud. Openpilot's biggest weakness has always been in the very "robotic" way that it drives…”
ramesh31 · Hacker News ↗
The benchmark is a curiosity rather than a meaningful test; specialized systems already outperform this proof-of-concept.
-
“Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model…”
prometheus1992 · Hacker News ↗
OpenAI Creator of GPT-6 Astrajyoung8607 openpilot external contributor, benchmark commenterAditya Ramabadr et al. Benchmark creatorsvaline HN commenter
How it unfolded 6 developments, newest first · click a bar or a number to jump articlespostscomments
-
6
Commenters debate whether benchmark is meaningful or a 'juggling bananas' test
Skeptics question the benchmark's utility, comparing it to arbitrary tasks like juggling bananas. Others suggest existing specialized models (Tesla's ~10-15B parameter system) already outperform this proof-of-concept. One commenter notes the cars took over five minutes to navigate the cone course.
“You can see this in the photos, it took over five minutes for the cars to get around the cone course.”
— odo1242, HN commenter · source -
This is what I first though of after I saw the drawing tool use demo. Right now we have specific AI that can generate images and drive cars, but tool use shows a process oriented AI that can use tools to drive a car, like we use our eyes/hands/feet. It's process, not task, that matters for AGI. And I think for artistic collaboration you just want…
2 more of the top 3 · 8 posts in this stretch
-
> Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.That and also the fact that (in spite of their usefulness) LLMs still so often do incredibly dumb shit without thinking of the consequences that the idea of having them drive in public is absurd.Recently was using claude code/opus 5 to diagnose an…
-
I wonder if this could solve a driving problem I have. I want an automated system to slowly drive the cars from the entrance of my neighborhood to their designated parking spaces. Right now the humans do this and they go too fast, and ignore the stop signs. I think it would be safer if all cars are automatically parked instead. It seems doable…
-
-
5
Benchmark creators respond to latency criticism, describe methodology
The researchers clarify that the benchmark is designed to measure frontier LLMs' out-of-the-box driving capability, not to demonstrate a practical autonomous system. They note the cars operate at extremely low speeds for safety and acknowledge latency as the core limitation.
“Great point! Yeah latency was one of the biggest issues here. To cope with that (and for safety reasons) the cars are driving at extremely low speeds.”
— aditya-ramabadr -
Great point! Yeah latency was one of the biggest issues here. To cope with that (and for safety reasons) the cars are driving at extremely low speeds. They also get timestamps with every tool call output etc so they can, in theory, "in context learn" about their own latency and choose motion durations and control how fast their iteration loop is…
2 more of the top 3 · 3 posts in this stretch
-
> Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output actuators? Looks like yes. Sure.Well, if the massive cloud models that are generalized and have a world model that's good enough, you can just distill them into smaller models. As a point of reference, the current gen of…
-
Looking forward to the juggling bananas benchmark. If Claude can only manage 5 and Astra does 6, clearly they have a better model.
-
-
4
Commenters propose hybrid architecture combining local control with cloud reasoning
A user suggests a synthesis where hardware handles low-level control and safety, while cloud LLMs handle high-level decision-making like whether to pass a car or adjust speed for traffic and weather—addressing openpilot's 'robotic' driving style.
“Eyes, control, and safety critical features on the hardware, higher level decision making to the cloud.”
— ramesh31 -
I'm not an expert in the LLM space, but I'm an external contributor to comma.ai's openpilot project and I'm and quite familiar with how its controls work, so I looked from that perspective. There's two questions here:1) Could a cloud-delivered LLM figure out how to drive this route, based on those input data and given access to those output…
2 more of the top 3 · 4 posts in this stretch
-
Perhaps there's a synthesis to be had though. Eyes, control, and safety critical features on the hardware, higher level decision making to the cloud. Openpilot's biggest weakness has always been in the very "robotic" way that it drives, which is technically correct but causes frustration for other drivers. Deciding "should I pass this car" is a…
-
As an external contributor to comma.ai, do you feel like you helped contribute to these deaths mentioned in the article?
-
-
3
Researchers acknowledge latency as the core barrier to real-world deployment
An openpilot contributor notes that while Astra can theoretically drive the course, cloud latency makes it impractical for real-world use. The benchmark operates at extremely low speeds and includes timestamp data to let the model learn its own latency in context, but this is framed as a proof-of-concept, not a viable system.
“Three reasons: latency, latency, and latency. openpilot's driving model updates the target curvature and acceleration at 20Hz.”
— jyoung8607 -
This also explains why Astra is so good at video generation. I have an Astra+Higgsfield setup. I could point it to a Github repo and ask it to generate a product walkthrough and it did a very good job by generating fake screens (e.g. with data filled in) from real ones - which wasn't possible in earlier models
-
-
2
Commenters attribute Astra's performance to superior vision and spatial reasoning
Contributors note Astra's strong performance on hard vision and spatial benchmarks—including ARC-3 scores and real-world game playing (Portal, Factorio, RimWorld)—and suggest this capability translates to driving tasks.
“What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke.”
— famouswaffles -
What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke. That huge gap between Astra and Fable (in this case) is basically every hard vison/spatial benchmark i've seen including non-benchmarks like playing games (Portal, Factorio, RimWorld).SpatialBench -…
2 more of the top 3 · 5 posts in this stretch
-
The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.It’s mostly a latency problem at this point. The models are too big to run locally, but…
-
I’m morbidly curious whether the (supposedly) superior compaction support in recent GPT models with an appropriate harness has anything to do with this. A conventional LLM with conventional attention is, of course, wildly unsuitable to continuous tasks like driving, but maybe as the technology advances it will improve in its ability to sort-of…
-
-
1
Tech community discusses implications for autonomous driving architecture
Hacker News commenters debate whether the benchmark signals a fundamental shift in autonomous driving, from specialized vision stacks and lane grammars toward end-to-end LLM control. Some frame it as validating generalist models over domain-specific systems; others emphasize practical barriers.
“The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it's maybe all about to give way to a single GPT looking at camera feeds…”
— valine -
background
Researchers publish drivingbench benchmark showing GPT-6 Astra driving a real car — A benchmark website goes live allowing frontier AI models to control a Toyota Corolla's steering, accelerator, and brakes on a fixed cone course, with up to 3 attempts per model. The site tests whether cloud-delivered LLMs can drive real vehicles and evaluate performance via traces and video.
Also covered reported alongside — the timeline has no entry for these yet
-
first by HN Best, 14h ago · also HN Frontpage
What people are saying 6 voices from 1 site · best of 21 · verbatim
- Yesterday
-
Looking forward to the inevitable "Astra can land a plane now, with no autopilot"
-
So a Taalas chip can run Llama 3.1 8B at 17000 TPS...does that mean if we could get Astra at similar speeds we could get self-driving for free?
-
There's also token RTT on top of network latency.. but what if you had a model running at 10k tps (like taalas' llama3b-8
-
Has anyone else noticed human drivers becoming more aggressive and causing more accidents than ever?I thought it was because my smaller town was overrun after COVID by transplants, but I'm hearing similar complaints from other places I was considering relocating to.Perhaps the solution will be robocars where, if there's a potential road rage…
-
> otherwise you can't reactI'm far from neuroscience, but humans don't need to operate at 20Hz to drive a car. And human reaction latency (event to measurable action) is often over 1s (under 1Hz).
-
Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).