Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
- ID
- 22775
- Status
- summarized
- Published
- 08 Sep 2026, 10:49 PM
- Fetched
- 10 Sep 2026, 12:03 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.5
- Created
- 10 Sep 2026, 12:07 AM
- Tags
- Audience
- developersvibe_codersai_ml_learners
What happened
Piotr Migdał benchmarked Qwen3.8 27B GGUF quantizations from Unsloth, spending ~$3,000 on Modal GPUs. The 4-bit Q4_K_M quantization (17 GB) matches the full BF16 model (55 GB) on Terminal-Bench 2.1 and fits on a 24 GB RTX 4090 with ~64k tokens of context, while 1-bit (6.2 GB) collapses to near random chance on GPQA Diamond.
Why it matters
If you run local LLMs for coding or agentic tasks, use Q4_K_M 4-bit quantization for Qwen3.8 27B — it fits on a single 24 GB consumer GPU with room for substantial context and loses no measurable benchmark quality. Avoid 1-bit and 2-bit quantizations for any reasoning workload, as longer reasoning chains make the degradation worse.
Discussion angle
The practical cutoff for quantization: 4-bit is free quality-wise, but where exactly does the cliff fall between 4-bit and 1-bit, and does it differ by task type (coding vs. science reasoning)?