AI Weekly Malaysia

Back to items Summaries

Transformers now runs llama.cpp quants

ID
27173
Status
summarized
Published
22 Sep 2026, 8:00 AM
Fetched
22 Sep 2026, 7:14 PM
Provider
Hugging Face Blog
Category
developer-ai
Original URL
https://huggingface.co/blog/transformers-llama-cpp-quants
Source URL
https://huggingface.co/blog/feed.xml

Summary

Score
6.5
Created
22 Sep 2026, 7:15 PM
Tags
Audience
developersai_ml_learnersvibe_coders

What happened

Hugging Face's transformers library now supports loading GGUF (llama.cpp quantized) models directly via from_pretrained, reusing llama.cpp's ggml kernels for performance. Initial support targets Apple Silicon and the Qwen3.5 architecture, with recommended starting quantization Q4_K_M (2.74 GB for Qwen3.5-4B vs 8.42 GB unquantized).

Why it matters

If you already build with transformers APIs, you can now run quantized local models without adding Ollama or LM Studio as a separate toolchain — but only on Apple Silicon and only for Qwen3.5-family models so far. Start with Q4_K_M, then move up to Q5_K_M or Q6_K if you have headroom; evaluate quality on your actual workload rather than assuming the smallest quant is fine.

Discussion angle

Is this enough to skip Ollama for local prototyping, or does the Apple-Silicon-only and Qwen3.5-only limitation make it a non-starter for anyone on Linux/Windows or using Llama/Mistral architectures today?

Top