AI Weekly Malaysia

Back to items Summaries

Qwen3.8 27B addition in words

ID
31940
Status
summarized
Published
05 Oct 2026, 7:34 AM
Fetched
06 Oct 2026, 1:46 AM
Provider
Simon Willison
Category
developer-ai
Original URL
https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/
Source URL
https://simonwillison.net/atom/everything/

Summary

Score
7.0
Created
06 Oct 2026, 1:46 AM
Tags
Audience
developersai_ml_learnersai_agent_users

What happened

Simon Willison re-ran a two-year-old GPT-4o experiment (originally posted by Colin Frasier on Bluesky) on local hardware, testing whether `Qwen3.8-27B-Q4_K_M.gguf` on a DGX Spark could add positive integers and return exact results only in English words. With reasoning disabled across 5,070 cases it hit 23.57% numeric accuracy, falling from 97.04% on one-to-three-digit operands to 6.44% on ten-to-thirteen-digit operands, even though format compliance was 96.17%. A paired 169-case run with medium reasoning enabled got 167/169 correct one-shot, with visible carry-by-carry traces in the report.

Why it matters

If you deploy a local quantized model with reasoning turned off to save latency, this is a direct measurement of the cost: 23.57% accuracy on word-form arithmetic versus 167/169 with reasoning on, on the same 27B Q4_K_M weights. The more dangerous number is the 96.17% format compliance — the model still emits well-formed English answers when it is wrong, so validating output shape is not validating output correctness. Anyone piping local-model output into anything that acts on numbers should add a real correctness check, or leave reasoning enabled for those paths and budget the extra latency (Willison notes the reasoning run took much longer per pair, which is why he dropped from 30 samples per cell to one).

Discussion angle

Should reasoning be on by default for local models, or is the latency worth paying only on specific paths? Walk through the 96.17% format-compliance-versus-23.57%-accuracy gap and ask what your own structured-output pipeline would do with a confidently formatted wrong number.

Top