GPT-6 Astra has gained the ability to drive a car
- ID
- 27934
- Status
- summarized
- Published
- 23 Sep 2026, 11:14 PM
- Fetched
- 24 Sep 2026, 7:00 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://drivingbench.com/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 6.5
- Created
- 24 Sep 2026, 7:00 AM
- Tags
- Audience
- developersai_agent_usersai_ml_learnersvibe_coders
What happened
DrivingBench hands frontier models control of a real Toyota Corolla's steering, accelerator and brakes on a fixed cone course, publishing per-attempt traces, videos, token counts and costs. GPT-6 Astra (Codex, medium) was the only model to finish: 100% course progress in 5:22 on its second attempt in the same chat, after attempt one stalled at 49% (DNF). Claude Fable 5.1 (Claude Code) peaked at 45%, Grok 4.6 (Cursor) 11% and GPT-5.6 Sol (Codex) 6%, all DNF; the Hacker News thread drew 258 points and 219 comments.
Why it matters
The leaderboard ranks best-of-up-to-3 attempts inside one continuous chat and shows first-attempt results beside it, so the single success is a retry after the model reflected on its own failure — and that winning run burned 246.6M tokens / $7.74, while Grok 4.6 spent $0.29 across three failed attempts. If you run or buy agent evals, report first-attempt success and best-of-3 separately, and price the retry loop rather than the single call, because tokens per attempt vary by roughly 100x across these models. No Malaysia-specific detail appears in the source text, so treat this as an eval-methodology item, not a local one.
Discussion angle
The only finished run needed a second attempt in the same context after a failed first one — so how much of your agent's measured reliability comes from same-context reflection, and would you be able to tell first-attempt success from best-of-3 in your own stack today?