AI Weekly Malaysia

Back to items Summaries

Jeeves. Reasoning improves Jev-like decision models

ID
30475
Status
summarized
Published
29 Sep 2026, 7:13 PM
Fetched
01 Oct 2026, 3:20 AM
Provider
Hacker News
Category
dev-community
Original URL
https://github.com/PostHog/jeeves
Source URL
https://hnrss.org/best

Summary

Score
7.0
Created
01 Oct 2026, 3:22 AM
Tags
Audience
developersai_ml_learnersai_agent_userssaas_founders

What happened

PostHog published Jeeves, an open-source reasoning classifier built on Qwen3.5-9B with LoRA plus a pointer head, trained with SFT and CISPO, and shipped with full training code and train/dev/test data. It reports 0.889 accuracy on held-out out-of-domain test data (vs Kev-9B 0.822 and Jev 0.857) and 0.935 on JevBench's 231 public items (vs Jev 0.866), using a block-4 diffusion drafter and a Jev-compatible API supporting noul/choice/score questions. Latency is about 0.3 s per request without thinking and a 3.3 s median with thinking on a single H100 at --precision fp8; it runs on CUDA bf16, FP8 on compute capability 8.9+, and Apple Silicon MPS. The thread drew 239 points and 93 comments on Hacker News.

Why it matters

The headline numbers hide a regression: Jeeves scores 0.746 on Transfer (MMLU-Pro and buried state) versus Jev's 0.800, and 0.793 on MMLU versus Jev's 0.900, so reasoning-before-deciding helps on the benchmarks it targets and hurts on general transfer tasks. If you currently fall back to a reasoning model when a Jev-like classifier is uncertain, the 3.3 s median thinking latency versus 0.3 s without means that fallback costs roughly an order of magnitude more wall-clock per request on one H100 at fp8 — decide per pipeline whether you truncate the chain, or keep the calibrated classifier and only reason on the hard slice. Because inference also runs on Apple Silicon in bf16 or FP8, you can benchmark it on a local Mac before paying for cloud GPU time.

Discussion angle

Jeeves beats Jev on JevBench (0.935 vs 0.866) but loses on Transfer (0.746 vs 0.800) and MMLU (0.793 vs 0.900) — is 'reason before you decide' worth 10x latency if it trades away general-task accuracy, and which of your own eval sets would catch that tradeoff?

Top