Getting the most out of Opus 5.5 in Claude and Claude Code
- ID
- 31573
- Status
- summarized
- Published
- 04 Oct 2026, 2:29 AM
- Fetched
- 04 Oct 2026, 12:44 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://claude.dev/blog/getting-the-most-out-of-opus-5-5/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 04 Oct 2026, 12:44 PM
- Tags
- Audience
- developersvibe_codersai_agent_users
What happened
A claude.dev blog guide by Addy Osmani (published Sep 22, 2026, 9 min read) walks through prompting Opus 5.5 in Claude apps and Claude Code, and the HN thread drew 191 points and 132 comments. Its concrete claims: Opus 5.5 always thinks before replying and decides how much, so "think carefully" / "think step by step" lines should be deleted from prompts and saved instructions — in the author's chat-product testing, removing one made replies start sooner with no clear quality drop. It also advises giving the whole task in one message with an explicit finish line (e.g. "the test suite passes", "every endpoint uses the new client") plus a stop-and-ask condition, and notes early testers had it run long coding tasks for hours with little oversight; in Claude Code, thinking depth is changed via an "effort" setting. The article text is truncated after section 2, so guidance on checking results, Claude apps, flagged messages, and speed is not available here.
Why it matters
If your saved prompts, CLAUDE.md, or agent system instructions still contain "think step by step" boilerplate, this says you can delete it and get faster first tokens with no measured quality loss — a one-line edit you can A/B this week. The bigger operational point: because the model runs for hours unsupervised on multi-step work, your prompt now needs a machine-checkable definition of done and an explicit stop condition, otherwise you are paying for and reviewing runs with no defined endpoint. Caveat: these are the author's own tests on a vendor-adjacent blog, not independent benchmarks — treat the latency claim as a hypothesis to verify on your own tasks.
Discussion angle
Grep your own prompts and instruction files for "think step by step" / "think carefully" — remove them and compare time-to-first-token and output quality on one real task. Then debate: if the model decides its own thinking budget, does prompt engineering shift from incantations to writing explicit finish lines and stop conditions, and how do you write a 'done' definition an agent can actually check?