Commenters debate whether benchmark is meaningful or a 'juggling bananas' test
6 Today 12:51 PM · 10h ago · 1 post · 2 comments · 2 sources · development 6 of 6
Skeptics question the benchmark's utility, comparing it to arbitrary tasks like juggling bananas. Others suggest existing specialized models (Tesla's ~10-15B parameter system) already outperform this proof-of-concept. One commenter notes the cars took over five minutes to navigate the cone course.
“You can see this in the photos, it took over five minutes for the cars to get around the cone course.”
odo1242, HN commenter · hn ↗OpenAI Creator of GPT-6 Astrajyoung8607 openpilot external contributor, benchmark commenterAditya Ramabadr et al. Benchmark creatorsvaline HN commenter
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 8 voices · verbatim
-
This is what I first though of after I saw the drawing tool use demo. Right now we have specific AI that can generate images and drive cars, but tool use shows a process oriented AI that can use tools to drive a car, like we use our eyes/hands/feet. It's process, not task, that matters for AGI. And I think for artistic collaboration you just want…
-
> Could this work in the real world? Absolutely not. Three reasons: latency, latency, and latency.That and also the fact that (in spite of their usefulness) LLMs still so often do incredibly dumb shit without thinking of the consequences that the idea of having them drive in public is absurd.Recently was using claude code/opus 5 to diagnose an…
-
I wonder if this could solve a driving problem I have. I want an automated system to slowly drive the cars from the entrance of my neighborhood to their designated parking spaces. Right now the humans do this and they go too fast, and ignore the stop signs. I think it would be safer if all cars are automatically parked instead. It seems doable…
-
Has anyone else noticed human drivers becoming more aggressive and causing more accidents than ever?I thought it was because my smaller town was overrun after COVID by transplants, but I'm hearing similar complaints from other places I was considering relocating to.Perhaps the solution will be robocars where, if there's a potential road rage…
-
So a Taalas chip can run Llama 3.1 8B at 17000 TPS...does that mean if we could get Astra at similar speeds we could get self-driving for free?
-
There's also token RTT on top of network latency.. but what if you had a model running at 10k tps (like taalas' llama3b-8
-
You can see this in the photos, it took over five minutes for the cars to get around the cone course.
-
Looking forward to the inevitable "Astra can land a plane now, with no autopilot"
All 6 developments of GPT-6 Astra successfully drives a real car in benchmark test →
Hacker NewsNewswiresMastodon