Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?
- ID
- 27491
- Status
- summarized
- Published
- 23 Sep 2026, 7:12 AM
- Fetched
- 23 Sep 2026, 8:18 AM
- Provider
- Lenny's Newsletter
- Category
- product-startup
- Original URL
- https://www.lennysnewsletter.com/p/opus-55-vs-gpt-6-sol-which-model
- Source URL
- https://www.lennysnewsletter.com/feed
Summary
- Score
- 7.5
- Created
- 23 Sep 2026, 8:18 AM
- Tags
- Audience
- developersvibe_codersai_agent_usersai_ml_learners
What happened
Claire Vo ran a blind evaluation of Claude Opus 5.5, GPT-6 Sol, GPT-6 Astra, and Luna across emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. Opus 5.5 came out strongest overall—especially for long-running agents and B2B frontend—while Sol won on clear writing, readable PRDs, and price, and Astra excelled at creative tasks. An LLM judge disagreed with her human rankings, and Barbie Bench (a 3D fashion-game test) showed hands are still 'tragic' and AGI hasn't arrived.
Why it matters
If you're picking a model for a specific workflow, this points to concrete splits: use Opus 5.5 for long-running agents and B2B frontend prototyping, Sol for PRDs and cost-sensitive writing tasks, and Astra for creative output. The LLM judge disagreement is a practical warning that automated eval pipelines may reward different qualities than your own taste—worth knowing before you build evals on top of these models.
Discussion angle
The LLM judge disagreed with the human's blind rankings—what was it rewarding that the human wasn't, and does that mean our automated eval harnesses are measuring the wrong things?