AI Weekly Malaysia

Back to items Summaries

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

ID
23929
Status
summarized
Published
13 Sep 2026, 4:25 AM
Fetched
15 Sep 2026, 4:45 AM
Provider
Hacker News
Category
dev-community
Original URL
https://withspecific.com/benchmarks/real-swe
Source URL
https://hnrss.org/best

Summary

Score
8.0
Created
15 Sep 2026, 5:51 AM
Tags
Audience
developersvibe_codersai_ml_learnersai_agent_users

What happened

Specific Labs released Real-SWE, a benchmark that tests frontier AI coding agents on private, licensed enterprise codebases with real business-critical tasks like billing, tax calculations, and customer migrations. The top model, Fable 5.1 via Claude Code, resolved only 38.8% of tasks (pass@1 averaged over 8 runs), with GPT-5.6 Sol via Codex CLI at the bottom at 16.2%. The benchmark deliberately uses proprietary code and company-specific conventions not available on the public internet.

Why it matters

If you are deciding which coding agent to deploy in a real production codebase, these numbers set a realistic expectation: even the best agent fails on roughly 6 out of 10 enterprise tasks, and performance varies widely by harness (e.g., GLM 5.3 on Claude Code beats Grok 4.6 on Grok Build). Do not extrapolate from synthetic or public-repo benchmarks when budgeting for agent-assisted engineering work on proprietary systems.

Discussion angle

Compare Real-SWE's sub-40% top resolution rate against the benchmarks you currently use to justify agent adoption—what does that gap mean for ROI when agents are handling billing or tax logic where failures have direct business cost?

Top