Summaries
Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.
Showing 1-3 of 3 results
| Date | Provider | Score | Summary |
|---|---|---|---|
| 29 Sep 2026, 4:23 AM | Hacker News | 7.8 | Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Jeff is an independent open-source project offering fine-tunes of Qwen3.5 (0.8B and 2B) and Gemma 4 (E2B) as tiny zero-shot classification models that reuse Jev's request format and return a calibrated probability per option from a single forward pass instead of generated text. The README reports about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max via MLX, with the 0.8B training in roughly 2 hours and the 2B in about 3.5 hours on one RTX PRO 6000, using synthetic data written by an open model on two DGX Sparks. It is explicitly not affiliated with or endorsed by TypeSafe, the makers of Jev, and the repo shows 298 stars, 8 forks and 6 commits; the Hacker News thread drew 222 points and 71 comments. Why: If you currently route simple label decisions — support queues, moderation labels, intents, game moves — through a hosted LLM API, this is a concrete alternative: ~22-28 ms per decision on a single GPU or an M4 Max MacBook, no per-token billing and no data leaving the machine. The reported fine-tune result (held-out accuracy 31.7% to 95.8% for voice navigation in under 30 minutes on one GPU) is the number to test against your own labels, since zero-shot accuracy at 0.8B is the stated weak point and the README itself says reasoning will not match a much larger model. For teams in Malaysia, running this on local or consumer hardware removes cloud GPU spend and cross-border data transfer for classification tasks, though you still need to verify the models' licensing and Jev's own terms before swapping them in. |
| 28 Sep 2026, 11:03 PM | Lenny's Newsletter | 7.0 | 🎙️ How I AI: Jev for beginners + I left Claude for months, Opus 5.5 brought me back + Opus 5.5 vs. GPT-6 Sol bench
In a solo 'How I AI' episode, Claire tests Jev, TypeSafe AI's decision model that returns structured values (categories, scores, probabilities) instead of generated text, and reports concrete costs: 9 cents to compare 1,700 ChatPRD pull requests across 17,000 pairs, 4,500 YouTube comments searched, and 200,000 classifications run for about $4. Pricing is stated as 4 cents per million input tokens with no output-token fee, and she pairs Jev with a frontier model for deeper reasoning on filtered subsets. She also notes Claude Code and Codex keep past sessions locally, and that her engineering usage fell from nearly 100% of her AI usage in January to under 40% by September. The excerpt covers only the Jev segment; the Opus 5.5 and GPT-6 Sol benchmark items named in the title are not detailed in the text provided. Why: If a chunk of your pipeline is classification, tagging, routing, or scoring, this is a concrete re-costing prompt: 4 cents per million input tokens with no output-token charge and a claimed ~$4 for 200,000 operations means workloads you previously considered too expensive at scale may now be worth building. The second actionable detail is local session history — Claude Code and Codex store past sessions on disk, so you can classify your own logs before committing to any new tooling. Treat the pricing and benchmarks as vendor-side claims from a single user's week, not independent measurement. |
| 28 Sep 2026, 8:03 PM | Lenny's Newsletter | 6.5 | Jev for beginners: how to use it and what to build
Claire Vo walks through Jev, TypeSafe AI's "decision model" that returns type-safe structured values (a choice, a score, a probability) instead of generated text, priced at 4 cents per million input tokens with no output charge. She reports running it on five projects in a week: categorizing 1,700 PRs for 9 cents, a meta-analysis of her own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph (1,100 signals, 200,000 classifications), and a live dashboard built from 4,500 YouTube comments. She also says she stopped using Jev alone and now pairs it with other models such as Gemini 3.5 Flash-Lite. Why: If your pipeline spends money on an LLM just to bucket, label, or score things, this is a concrete alternative pricing shape to test: input-only billing with no output charge, claimed at 4 cents per million input tokens and 9 cents for 1,700 PR categorizations. The practical move is to take one existing classification or triage job you already run and benchmark a structured-output decision model against your current model on cost and label accuracy, rather than assuming general chat-model pricing. Note this is a launch-week episode with a sponsor segment, so the numbers are the author's own reported results, not an independent benchmark, and there is no Malaysia or Southeast Asia angle in the text. |