Advanced evals: How to find (and fix) hidden AI failures in your product
- ID
- 27234
- Status
- summarized
- Published
- 22 Sep 2026, 8:45 PM
- Fetched
- 22 Sep 2026, 9:23 PM
- Provider
- Lenny's Newsletter
- Category
- product-startup
- Original URL
- https://www.lennysnewsletter.com/p/advanced-evals-how-to-find-and-fix
- Source URL
- https://www.lennysnewsletter.com/feed
Summary
- Score
- 8.0
- Created
- 22 Sep 2026, 9:24 PM
- Tags
- Audience
- developersai_ml_learnerssaas_foundersai_agent_users
What happened
Hamel Husain and Shreya Shankar argue most AI teams skip error discovery and jump straight to writing metrics, measuring the wrong things. Drawing from work with 50+ AI companies, they share which eval steps can and cannot be automated, plus a free plugin that lets a coding agent handle much of the eval heavy lifting. Concrete case studies include Shopify's eval-guided AI workflow builder (2.2x faster, 68% cheaper than the frontier-model system it replaced), Cursor's Auto Balance routing (41% cost reduction), and Ramp's receipt collection accuracy jumping from 35% to 83%.
Why it matters
If you ship AI features, you should stop jumping to metric definitions and invest in error discovery first—the step most teams skip. The Shopify and Cursor examples show evals can cut costs 40-68% while improving quality, so this is a direct margin and UX lever, not just a QA exercise. The free coding-agent plugin means even small teams can start without a dedicated evals engineer.
Discussion angle
Compare your current eval process against the 'error discovery first, metrics second' framing—do you have a concrete error-discovery step, or are you already writing metrics blind? Share one eval gap you found recently and whether a coding-agent plugin could have surfaced it faster.