AI Weekly Malaysia

Back to items Summaries

GPT-6 Astra on robot arms

ID
21961
Status
summarized
Published
06 Sep 2026, 9:52 AM
Fetched
08 Sep 2026, 5:23 AM
Provider
Hacker News
Category
dev-community
Original URL
https://openai.robocurve.org/gpt-6-astra/
Source URL
https://hnrss.org/best

Summary

Score
7.0
Created
08 Sep 2026, 5:24 AM
Tags
Audience
developersai_agent_usersai_ml_learners

What happened

Robocurve benchmarked GPT-6 Astra against Claude Fable 5 and 5.1 on identical YAM robot arm tasks. Astra completed the block-into-bowl task 19/20 times (95%) at $0.94/run in 2.5 min, versus Fable 5.1's 8/20 (40%) at $2.12/run in 6.8 min. On the puzzle-insertion task, both models stalled at the same final step, completing only 2/20 each.

Why it matters

If you're evaluating LLM-driven agents for physical manipulation, Astra is roughly 2x cheaper and 3x faster on coarse pick-and-place, but precision insertion remains unsolved across both model families—budget for human intervention or classical control on that last-mile step rather than expecting any frontier model to close the gap.

Discussion angle

Both frontier models fail identically on the puzzle insertion—does this suggest a shared architectural limitation in how vision-language models handle fine-grained force feedback, and would a hybrid approach (LLM planning + classical controller for insertion) be more cost-effective than waiting for GPT-7?

Top