AI Weekly Malaysia

Back to items Summaries

GPT-6 Astra has gained the ability to drive a car

ID
27934
Status
summarized
Published
23 Sep 2026, 11:14 PM
Fetched
24 Sep 2026, 7:00 AM
Provider
Hacker News
Category
dev-community
Original URL
https://drivingbench.com/
Source URL
https://hnrss.org/best

Summary

Score
6.5
Created
24 Sep 2026, 7:00 AM
Tags
Audience
developersai_agent_usersai_ml_learnersvibe_coders

What happened

DrivingBench hands frontier models control of a real Toyota Corolla's steering, accelerator and brakes on a fixed cone course, publishing per-attempt traces, videos, token counts and costs. GPT-6 Astra (Codex, medium) was the only model to finish: 100% course progress in 5:22 on its second attempt in the same chat, after attempt one stalled at 49% (DNF). Claude Fable 5.1 (Claude Code) peaked at 45%, Grok 4.6 (Cursor) 11% and GPT-5.6 Sol (Codex) 6%, all DNF; the Hacker News thread drew 258 points and 219 comments.

Why it matters

The leaderboard ranks best-of-up-to-3 attempts inside one continuous chat and shows first-attempt results beside it, so the single success is a retry after the model reflected on its own failure — and that winning run burned 246.6M tokens / $7.74, while Grok 4.6 spent $0.29 across three failed attempts. If you run or buy agent evals, report first-attempt success and best-of-3 separately, and price the retry loop rather than the single call, because tokens per attempt vary by roughly 100x across these models. No Malaysia-specific detail appears in the source text, so treat this as an eval-methodology item, not a local one.

Discussion angle

The only finished run needed a second attempt in the same context after a failed first one — so how much of your agent's measured reliability comes from same-context reflection, and would you be able to tell first-attempt success from best-of-3 in your own stack today?

Top