AI Weekly Malaysia

Back to items Summaries

The Normalization of Inexplicable Failures

ID
29132
Status
summarized
Published
27 Sep 2026, 11:26 PM
Fetched
28 Sep 2026, 1:59 AM
Provider
Hacker News
Category
dev-community
Original URL
https://www.ihatethefuture.com/2026/09/the-normalization-of-inexplicable.html
Source URL
https://hnrss.org/best

Summary

Score
7.0
Created
28 Sep 2026, 2:00 AM
Tags
Audience
developersai_ml_learnersai_agent_usersfounders

What happened

A blog post on ihatethefuture.com argues that Jev, an AI model from TypeSafe AI that returns typed values with probability estimates, sells speed and cost while leaving the hard part — evals and a ground-truth pipeline — to the buyer. The author's core complaint is that confidence scores are useless without knowing how well calibrated they are and without a model of what each wrong answer costs; Jev's marketing leans on benchmark scores, not calibration. The piece points out that Jev's own docs invent thresholds — 0.5 for 'do nothing' and 0.9 for 'do high-risk actions' — with no stated basis. The Hacker News thread drew 150 points and 41 comments.

Why it matters

If you ship an AI feature that branches on a confidence score, this is the concrete failure mode: you inherit a number you have never measured the calibration of, and you set a threshold by vibe (the article cites 0.5 and 0.9 defaults appearing in the vendor's own docs). Decide now whether you can produce a ground-truth set and a cost model per wrong answer — the author's claim is that if you can, you are already most of the way to fine-tuning your own solution instead of buying the model. If you cannot, treat the confidence score as decoration and design the product so a wrong answer is cheap.

Discussion angle

Ask everyone who has shipped a confidence threshold what number they picked and how they validated it — then check whether anyone has measured calibration on their own traffic rather than reading a benchmark table. The article's 0.5/'do nothing' and 0.9/'high-risk action' defaults are a good live example of a threshold chosen because it sounded right.

Top