AI Weekly Malaysia

Back to items Summaries

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

ID
27491
Status
summarized
Published
23 Sep 2026, 7:12 AM
Fetched
23 Sep 2026, 8:18 AM
Provider
Lenny's Newsletter
Category
product-startup
Original URL
https://www.lennysnewsletter.com/p/opus-55-vs-gpt-6-sol-which-model
Source URL
https://www.lennysnewsletter.com/feed

Summary

Score
7.5
Created
23 Sep 2026, 8:18 AM
Tags
Audience
developersvibe_codersai_agent_usersai_ml_learners

What happened

Claire Vo ran a blind evaluation of Claude Opus 5.5, GPT-6 Sol, GPT-6 Astra, and Luna across emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. Opus 5.5 came out strongest overall—especially for long-running agents and B2B frontend—while Sol won on clear writing, readable PRDs, and price, and Astra excelled at creative tasks. An LLM judge disagreed with her human rankings, and Barbie Bench (a 3D fashion-game test) showed hands are still 'tragic' and AGI hasn't arrived.

Why it matters

If you're picking a model for a specific workflow, this points to concrete splits: use Opus 5.5 for long-running agents and B2B frontend prototyping, Sol for PRDs and cost-sensitive writing tasks, and Astra for creative output. The LLM judge disagreement is a practical warning that automated eval pipelines may reward different qualities than your own taste—worth knowing before you build evals on top of these models.

Discussion angle

The LLM judge disagreed with the human's blind rankings—what was it rewarding that the human wasn't, and does that mean our automated eval harnesses are measuring the wrong things?

Top