Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
- ID
- 23929
- Status
- summarized
- Published
- 13 Sep 2026, 4:25 AM
- Fetched
- 15 Sep 2026, 4:45 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://withspecific.com/benchmarks/real-swe
- Source URL
- https://hnrss.org/best
Summary
- Score
- 8.0
- Created
- 15 Sep 2026, 5:51 AM
- Tags
- Audience
- developersvibe_codersai_ml_learnersai_agent_users
What happened
Specific Labs released Real-SWE, a benchmark that tests frontier AI coding agents on private, licensed enterprise codebases with real business-critical tasks like billing, tax calculations, and customer migrations. The top model, Fable 5.1 via Claude Code, resolved only 38.8% of tasks (pass@1 averaged over 8 runs), with GPT-5.6 Sol via Codex CLI at the bottom at 16.2%. The benchmark deliberately uses proprietary code and company-specific conventions not available on the public internet.
Why it matters
If you are deciding which coding agent to deploy in a real production codebase, these numbers set a realistic expectation: even the best agent fails on roughly 6 out of 10 enterprise tasks, and performance varies widely by harness (e.g., GLM 5.3 on Claude Code beats Grok 4.6 on Grok Build). Do not extrapolate from synthetic or public-repo benchmarks when budgeting for agent-assisted engineering work on proprietary systems.
Discussion angle
Compare Real-SWE's sub-40% top resolution rate against the benchmarks you currently use to justify agent adoption—what does that gap mean for ROI when agents are handling billing or tax logic where failures have direct business cost?