MicroLLM Lab – Try 7 tiny LLM's in the browser
- ID
- 29945
- Status
- summarized
- Published
- 29 Sep 2026, 2:58 AM
- Fetched
- 30 Sep 2026, 2:08 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://stateofutopia.com/experiments/microllmlab/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 30 Sep 2026, 2:10 AM
- Tags
- Audience
- developersai_ml_learnersvibe_codersai_agent_users
What happened
MicroLLM Lab is a browser-based playground that runs 7 tiny language models (25M–360M parameters) fully on-device using WebGPU and Q4 quantization, caching them in IndexedDB with no accounts or server. Q4 shrinks weights from 16-bit to 4 bits per parameter (a claimed 75% memory reduction), so 100M+ models fit in roughly 50–84 MB of browser memory. It ships a benchmark tab that scores speed (tokens/s sustained over a 256-token decode) and accuracy via objective regex/exact-token checks rather than writing quality, and offers a 589 MB zip download; the Hacker News thread drew 272 points and 111 comments.
Why it matters
If you currently pay per-token just to classify, filter spam, or extract intent before calling a frontier model, this gives you a free way to measure whether a 25M–360M Q4 model can do that triage step on the client instead — and its accuracy benchmark returns a pass-rate number on objective checks, not vibes. Two practical caveats from the page: the full model bundle is a 589 MB download and models cache into each user's browser IndexedDB, so bandwidth and first-load UX are real costs for your users. The claimed sub-10ms time-to-first-token is the author's own figure — verify it against your own hardware before designing a real-time autocomplete flow around it.
Discussion angle
Run the suite on every loaded model on your own laptop and compare the sustained tok/s and accuracy pass rates against the page's claims — then decide whether local SLM triage is good enough to replace a cloud classify call in one specific step of your agent or app pipeline.