How Warp ships 2,000 PRs a month with AI factories | Zach Lloyd (CEO, Warp)
- ID
- 26816
- Status
- summarized
- Published
- 21 Sep 2026, 8:04 PM
- Fetched
- 21 Sep 2026, 9:56 PM
- Provider
- Lenny's Newsletter
- Category
- product-startup
- Original URL
- https://www.lennysnewsletter.com/p/how-warp-ships-2000-prs-a-month-with
- Source URL
- https://www.lennysnewsletter.com/feed
Summary
- Score
- 7.0
- Created
- 21 Sep 2026, 9:57 PM
- Tags
- Audience
- developersai_agent_userssaas_founders
What happened
Warp CEO Zach Lloyd describes how his team uses a cloud-based AI software factory called 'Wilson' to ship 2,000 PRs per month, routing requests from Slack through Linear and GitHub into tested pull requests. He details concrete operational signals they track—human interactions per PR, cost per PR across model configs, and LLM-as-a-judge scoring of every agent run—and explains how the factory self-improves by replaying failed runs to build cost-quality Pareto charts for model selection.
Why it matters
If you are building or evaluating AI coding pipelines, the specific metrics Warp tracks (human interactions per PR, cost per PR, LLM-as-judge scores per run) are a usable blueprint for measuring your own agent throughput rather than guessing at quality. The admission that human review remains the bottleneck is a practical signal: invest in review tooling and review-agent workflows, not just generation agents.
Discussion angle
Which of Warp's metrics—human interactions per PR, cost per PR, or LLM-as-judge run scores—would actually work in your current team's tooling, and which would be noise?