Qwen3.8 27B addition in words
- ID
- 31940
- Status
- summarized
- Published
- 05 Oct 2026, 7:34 AM
- Fetched
- 06 Oct 2026, 1:46 AM
- Provider
- Simon Willison
- Category
- developer-ai
- Original URL
- https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/
- Source URL
- https://simonwillison.net/atom/everything/
Summary
- Score
- 7.0
- Created
- 06 Oct 2026, 1:46 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_users
What happened
Simon Willison re-ran a two-year-old GPT-4o experiment (originally posted by Colin Frasier on Bluesky) on local hardware, testing whether `Qwen3.8-27B-Q4_K_M.gguf` on a DGX Spark could add positive integers and return exact results only in English words. With reasoning disabled across 5,070 cases it hit 23.57% numeric accuracy, falling from 97.04% on one-to-three-digit operands to 6.44% on ten-to-thirteen-digit operands, even though format compliance was 96.17%. A paired 169-case run with medium reasoning enabled got 167/169 correct one-shot, with visible carry-by-carry traces in the report.
Why it matters
If you deploy a local quantized model with reasoning turned off to save latency, this is a direct measurement of the cost: 23.57% accuracy on word-form arithmetic versus 167/169 with reasoning on, on the same 27B Q4_K_M weights. The more dangerous number is the 96.17% format compliance — the model still emits well-formed English answers when it is wrong, so validating output shape is not validating output correctness. Anyone piping local-model output into anything that acts on numbers should add a real correctness check, or leave reasoning enabled for those paths and budget the extra latency (Willison notes the reasoning run took much longer per pair, which is why he dropped from 30 samples per cell to one).
Discussion angle
Should reasoning be on by default for local models, or is the latency worth paying only on specific paths? Walk through the 96.17% format-compliance-versus-23.57%-accuracy gap and ask what your own structured-output pipeline would do with a confidently formatted wrong number.