PrismML hopes its tiny LLM will change how we all use AI
- ID
- 25748
- Status
- summarized
- Published
- 18 Sep 2026, 6:34 AM
- Fetched
- 18 Sep 2026, 7:09 AM
- Provider
- TechCrunch
- Category
- technology
- Original URL
- https://techcrunch.com/2026/09/17/prismml-hopes-its-tiny-llm-could-change-how-we-all-use-ai/
- Source URL
- https://techcrunch.com/feed/
Summary
- Score
- 7.5
- Created
- 18 Sep 2026, 7:09 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_userssaas_founders
What happened
PrismML released Bonsai 2 27B, compressing Alibaba's Qwen3.8 27B down to 5.9 GB — a 9-10x memory reduction — while retaining 98% of the original's aggregate benchmark scores. The Caltech-founded startup (seed: $22.25M from Khosla Ventures, Cerberus, Caltech) targets on-device deployment on PCs and possibly high-end smartphones, and its first Bonsai model has been downloaded over 11 million times since March.
Why it matters
If you're paying for cloud GPU inference or API calls for 27B-class reasoning models, a 5.9 GB model that fits on a local PC could meaningfully cut your inference bill and latency. Builders in Malaysia and SEA where cloud GPU costs are high should test Bonsai 2 against their current Qwen3.8 workflows before committing to new infrastructure spend.
Discussion angle
Compare Bonsai 2's 98% benchmark retention claim against real-world task performance — benchmarks can mask degradation on specific workloads, so what should builders actually test before switching from Qwen3.8 to Bonsai 2?