HarnessTax: How Much Does the Harness Matter for Coding Agents?
- ID
- 25916
- Status
- summarized
- Published
- 17 Sep 2026, 6:10 AM
- Fetched
- 18 Sep 2026, 9:04 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://harnesstax.github.io/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 8.0
- Created
- 18 Sep 2026, 9:05 PM
- Tags
- Audience
- developersvibe_codersai_agent_users
What happened
UC Berkeley researchers evaluated 21 model–harness pairs across 7 models and 3 harnesses (Claude Code, Codex CLI, and Pi) on SWE-bench Lite and Terminal-Bench 2.0. They found that harness choice has little effect on task success rate but can cause up to 5x cost differences, that the minimal open-source harness Pi is competitive on both cost and success rate, and that models sometimes perform better with a different harness than their own vendor's.
Why it matters
If you're paying for Claude Code or Codex CLI, you may be spending up to 5x more for the same task success rate you'd get with a minimal open-source harness like Pi. Before committing to a vendor's harness, benchmark your actual workload across alternatives — the model matters more than the harness, and the harness mainly determines your cost.
Discussion angle
Try running the same coding task through Claude Code vs. Pi with the same model and compare both success rate and token cost live — the 5x cost gap is the kind of thing that changes tooling decisions immediately.