Sonnet 5.5
- ID
- 29498
- Status
- summarized
- Published
- 29 Sep 2026, 1:58 AM
- Fetched
- 29 Sep 2026, 4:59 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://www.anthropic.com/claude-sonnet-5-5
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.5
- Created
- 29 Sep 2026, 5:00 AM
- Tags
- Audience
- developersai_agent_usersvibe_codersai_ml_learners
What happened
Anthropic introduced Claude Sonnet 5.5, the second model in the Claude 5.5 family, claiming 30%+ faster output and up to 30% lower cost per task than Sonnet 5 at unchanged list pricing of $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads. It scores 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%, comes within two points of Opus 5.5 on GDPval-AA, and is the first Sonnet model to ship with cyber safeguards and fallbacks; Haiku 5.5 is promised in the coming weeks. The Hacker News thread drew 390 points and 254 comments.
Why it matters
If your coding agent or document pipeline defaults to Opus 5.5, this is a concrete reason to re-test model routing: Sonnet 5.5 claims 70.6% on Terminal-Bench 4.0 (the table lists Opus 5.5 at 66.4%, with a footnote) at $2/$10 per million tokens and 30%+ faster generation, so the cheaper model may now win on well-scoped bug fixes and slide/spreadsheet generation. Note these are Anthropic's own benchmark and cost figures — the 10.3% to 70.6% jump is large enough that you should run your own repo tasks through both before switching a default. Also flag the new cyber safeguards on a Sonnet-tier model: Anthropic says routine software development is unaffected, but anything security-adjacent you route through Sonnet may now hit fallbacks. For teams billing API usage in USD against MYR budgets, the token-efficiency claim (same per-token price, up to 30% fewer tokens per task) is the number to verify on your own workload.
Discussion angle
The Terminal-Bench 4.0 jump from 10.3% to 70.6% is the headline, but it's self-reported and the Opus 5.5 comparison sits in a footnoted table — what would you actually measure on your own tasks before swapping Sonnet 5.5 in as your default agent model, and does the Haiku 5.5 promise change how you'd tier your routing?