AI Weekly Malaysia

Back to items Summaries

HarnessTax: How Much Does the Harness Matter for Coding Agents?

ID
25916
Status
summarized
Published
17 Sep 2026, 6:10 AM
Fetched
18 Sep 2026, 9:04 PM
Provider
Hacker News
Category
dev-community
Original URL
https://harnesstax.github.io/
Source URL
https://hnrss.org/best

Summary

Score
8.0
Created
18 Sep 2026, 9:05 PM
Tags
Audience
developersvibe_codersai_agent_users

What happened

UC Berkeley researchers evaluated 21 model–harness pairs across 7 models and 3 harnesses (Claude Code, Codex CLI, and Pi) on SWE-bench Lite and Terminal-Bench 2.0. They found that harness choice has little effect on task success rate but can cause up to 5x cost differences, that the minimal open-source harness Pi is competitive on both cost and success rate, and that models sometimes perform better with a different harness than their own vendor's.

Why it matters

If you're paying for Claude Code or Codex CLI, you may be spending up to 5x more for the same task success rate you'd get with a minimal open-source harness like Pi. Before committing to a vendor's harness, benchmark your actual workload across alternatives — the model matters more than the harness, and the harness mainly determines your cost.

Discussion angle

Try running the same coding task through Claude Code vs. Pi with the same model and compare both success rate and token cost live — the 5x cost gap is the kind of thing that changes tooling decisions immediately.

Top