Can gzip be a language model?
- ID
- 27321
- Status
- summarized
- Published
- 22 Sep 2026, 2:08 PM
- Fetched
- 24 Sep 2026, 4:22 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://nathan.rs/posts/gzip-lm/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 5.5
- Created
- 24 Sep 2026, 4:25 AM
- Tags
- Audience
- developersai_ml_learners
What happened
Nathan Barry explores whether gzip can function as a language model by exploiting the compression-prediction equivalence from information theory. By priming gzip's 32 KiB DEFLATE sliding window with a corpus and scoring candidate continuations by their compressed length, he generates rough Shakespeare-like text via beam search. The output is incoherent but surprisingly structured, demonstrating that any compressor implicitly contains a probability model.
Why it matters
This is a teaching tool, not a shipping tool. It concretely demonstrates why compression and prediction are mathematically the same problem (bits = -log2 p), which is useful if you're learning information theory fundamentals or want to understand why LLMs can be framed as compressors. Don't expect to replace any real LM with gzip.
Discussion angle
The compression-prediction equivalence is increasingly cited in LLM scaling arguments—if you accept that better compression equals better prediction, then benchmarks like compression ratio become a model-agnostic way to compare LLMs. Worth discussing whether this framing changes how we think about model evaluation.