Jeeves. Reasoning improves Jev-like decision models
- ID
- 30475
- Status
- summarized
- Published
- 29 Sep 2026, 7:13 PM
- Fetched
- 01 Oct 2026, 3:20 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://github.com/PostHog/jeeves
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 01 Oct 2026, 3:22 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_userssaas_founders
What happened
PostHog published Jeeves, an open-source reasoning classifier built on Qwen3.5-9B with LoRA plus a pointer head, trained with SFT and CISPO, and shipped with full training code and train/dev/test data. It reports 0.889 accuracy on held-out out-of-domain test data (vs Kev-9B 0.822 and Jev 0.857) and 0.935 on JevBench's 231 public items (vs Jev 0.866), using a block-4 diffusion drafter and a Jev-compatible API supporting noul/choice/score questions. Latency is about 0.3 s per request without thinking and a 3.3 s median with thinking on a single H100 at --precision fp8; it runs on CUDA bf16, FP8 on compute capability 8.9+, and Apple Silicon MPS. The thread drew 239 points and 93 comments on Hacker News.
Why it matters
The headline numbers hide a regression: Jeeves scores 0.746 on Transfer (MMLU-Pro and buried state) versus Jev's 0.800, and 0.793 on MMLU versus Jev's 0.900, so reasoning-before-deciding helps on the benchmarks it targets and hurts on general transfer tasks. If you currently fall back to a reasoning model when a Jev-like classifier is uncertain, the 3.3 s median thinking latency versus 0.3 s without means that fallback costs roughly an order of magnitude more wall-clock per request on one H100 at fp8 — decide per pipeline whether you truncate the chain, or keep the calibrated classifier and only reason on the hard slice. Because inference also runs on Apple Silicon in bf16 or FP8, you can benchmark it on a local Mac before paying for cloud GPU time.
Discussion angle
Jeeves beats Jev on JevBench (0.935 vs 0.866) but loses on Transfer (0.746 vs 0.800) and MMLU (0.793 vs 0.900) — is 'reason before you decide' worth 10x latency if it trades away general-task accuracy, and which of your own eval sets would catch that tradeoff?