Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
- ID
- 29594
- Status
- summarized
- Published
- 29 Sep 2026, 4:23 AM
- Fetched
- 29 Sep 2026, 9:13 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://github.com/firelex/jeff
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.8
- Created
- 29 Sep 2026, 9:13 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coderssaas_founders
What happened
Jeff is an independent open-source project offering fine-tunes of Qwen3.5 (0.8B and 2B) and Gemma 4 (E2B) as tiny zero-shot classification models that reuse Jev's request format and return a calibrated probability per option from a single forward pass instead of generated text. The README reports about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max via MLX, with the 0.8B training in roughly 2 hours and the 2B in about 3.5 hours on one RTX PRO 6000, using synthetic data written by an open model on two DGX Sparks. It is explicitly not affiliated with or endorsed by TypeSafe, the makers of Jev, and the repo shows 298 stars, 8 forks and 6 commits; the Hacker News thread drew 222 points and 71 comments.
Why it matters
If you currently route simple label decisions — support queues, moderation labels, intents, game moves — through a hosted LLM API, this is a concrete alternative: ~22-28 ms per decision on a single GPU or an M4 Max MacBook, no per-token billing and no data leaving the machine. The reported fine-tune result (held-out accuracy 31.7% to 95.8% for voice navigation in under 30 minutes on one GPU) is the number to test against your own labels, since zero-shot accuracy at 0.8B is the stated weak point and the README itself says reasoning will not match a much larger model. For teams in Malaysia, running this on local or consumer hardware removes cloud GPU spend and cross-border data transfer for classification tasks, though you still need to verify the models' licensing and Jev's own terms before swapping them in.
Discussion angle
Where does a 22 ms probability-per-option call actually beat an LLM API call in your stack, and would you trust a 0.8B model for that decision or would the 31.7%-to-95.8% fine-tune path be mandatory before shipping?