Transformers now runs llama.cpp quants
- ID
- 27173
- Status
- summarized
- Published
- 22 Sep 2026, 8:00 AM
- Fetched
- 22 Sep 2026, 7:14 PM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/transformers-llama-cpp-quants
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.5
- Created
- 22 Sep 2026, 7:15 PM
- Tags
- Audience
- developersai_ml_learnersvibe_coders
What happened
Hugging Face's transformers library now supports loading GGUF (llama.cpp quantized) models directly via from_pretrained, reusing llama.cpp's ggml kernels for performance. Initial support targets Apple Silicon and the Qwen3.5 architecture, with recommended starting quantization Q4_K_M (2.74 GB for Qwen3.5-4B vs 8.42 GB unquantized).
Why it matters
If you already build with transformers APIs, you can now run quantized local models without adding Ollama or LM Studio as a separate toolchain — but only on Apple Silicon and only for Qwen3.5-family models so far. Start with Q4_K_M, then move up to Q5_K_M or Q6_K if you have headroom; evaluate quality on your actual workload rather than assuming the smallest quant is fine.
Discussion angle
Is this enough to skip Ollama for local prototyping, or does the Apple-Silicon-only and Qwen3.5-only limitation make it a non-starter for anyone on Linux/Windows or using Llama/Mistral architectures today?