🎙️ How I AI: Jev for beginners + I left Claude for months, Opus 5.5 brought me back + Opus 5.5 vs. GPT-6 Sol bench
- ID
- 29356
- Status
- summarized
- Published
- 28 Sep 2026, 11:03 PM
- Fetched
- 28 Sep 2026, 11:42 PM
- Provider
- Lenny's Newsletter
- Category
- product-startup
- Original URL
- https://www.lennysnewsletter.com/p/how-i-ai-jev-for-beginners-i-left
- Source URL
- https://www.lennysnewsletter.com/feed
Summary
- Score
- 7.0
- Created
- 28 Sep 2026, 11:42 PM
- Tags
- Audience
- developersvibe_codersai_ml_learnerssaas_founders
What happened
In a solo 'How I AI' episode, Claire tests Jev, TypeSafe AI's decision model that returns structured values (categories, scores, probabilities) instead of generated text, and reports concrete costs: 9 cents to compare 1,700 ChatPRD pull requests across 17,000 pairs, 4,500 YouTube comments searched, and 200,000 classifications run for about $4. Pricing is stated as 4 cents per million input tokens with no output-token fee, and she pairs Jev with a frontier model for deeper reasoning on filtered subsets. She also notes Claude Code and Codex keep past sessions locally, and that her engineering usage fell from nearly 100% of her AI usage in January to under 40% by September. The excerpt covers only the Jev segment; the Opus 5.5 and GPT-6 Sol benchmark items named in the title are not detailed in the text provided.
Why it matters
If a chunk of your pipeline is classification, tagging, routing, or scoring, this is a concrete re-costing prompt: 4 cents per million input tokens with no output-token charge and a claimed ~$4 for 200,000 operations means workloads you previously considered too expensive at scale may now be worth building. The second actionable detail is local session history — Claude Code and Codex store past sessions on disk, so you can classify your own logs before committing to any new tooling. Treat the pricing and benchmarks as vendor-side claims from a single user's week, not independent measurement.
Discussion angle
Is the 'cheap decision model does the filtering, frontier model does the reasoning' split worth adopting, or does it just add a second dependency? A practical test: classify your own local Claude Code or Codex session logs and compare the token spend against sending the same volume straight to a frontier model.