The Normalization of Inexplicable Failures
- ID
- 29132
- Status
- summarized
- Published
- 27 Sep 2026, 11:26 PM
- Fetched
- 28 Sep 2026, 1:59 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://www.ihatethefuture.com/2026/09/the-normalization-of-inexplicable.html
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 28 Sep 2026, 2:00 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersfounders
What happened
A blog post on ihatethefuture.com argues that Jev, an AI model from TypeSafe AI that returns typed values with probability estimates, sells speed and cost while leaving the hard part — evals and a ground-truth pipeline — to the buyer. The author's core complaint is that confidence scores are useless without knowing how well calibrated they are and without a model of what each wrong answer costs; Jev's marketing leans on benchmark scores, not calibration. The piece points out that Jev's own docs invent thresholds — 0.5 for 'do nothing' and 0.9 for 'do high-risk actions' — with no stated basis. The Hacker News thread drew 150 points and 41 comments.
Why it matters
If you ship an AI feature that branches on a confidence score, this is the concrete failure mode: you inherit a number you have never measured the calibration of, and you set a threshold by vibe (the article cites 0.5 and 0.9 defaults appearing in the vendor's own docs). Decide now whether you can produce a ground-truth set and a cost model per wrong answer — the author's claim is that if you can, you are already most of the way to fine-tuning your own solution instead of buying the model. If you cannot, treat the confidence score as decoration and design the product so a wrong answer is cheap.
Discussion angle
Ask everyone who has shipped a confidence threshold what number they picked and how they validated it — then check whether anyone has measured calibration on their own traffic rather than reading a benchmark table. The article's 0.5/'do nothing' and 0.9/'high-risk action' defaults are a good live example of a threshold chosen because it sounded right.