AI Weekly Malaysia

Back to items Summaries

MicroLLM Lab – Try 7 tiny LLM's in the browser

ID
29945
Status
summarized
Published
29 Sep 2026, 2:58 AM
Fetched
30 Sep 2026, 2:08 AM
Provider
Hacker News
Category
dev-community
Original URL
https://stateofutopia.com/experiments/microllmlab/
Source URL
https://hnrss.org/best

Summary

Score
7.0
Created
30 Sep 2026, 2:10 AM
Tags
Audience
developersai_ml_learnersvibe_codersai_agent_users

What happened

MicroLLM Lab is a browser-based playground that runs 7 tiny language models (25M–360M parameters) fully on-device using WebGPU and Q4 quantization, caching them in IndexedDB with no accounts or server. Q4 shrinks weights from 16-bit to 4 bits per parameter (a claimed 75% memory reduction), so 100M+ models fit in roughly 50–84 MB of browser memory. It ships a benchmark tab that scores speed (tokens/s sustained over a 256-token decode) and accuracy via objective regex/exact-token checks rather than writing quality, and offers a 589 MB zip download; the Hacker News thread drew 272 points and 111 comments.

Why it matters

If you currently pay per-token just to classify, filter spam, or extract intent before calling a frontier model, this gives you a free way to measure whether a 25M–360M Q4 model can do that triage step on the client instead — and its accuracy benchmark returns a pass-rate number on objective checks, not vibes. Two practical caveats from the page: the full model bundle is a 589 MB download and models cache into each user's browser IndexedDB, so bandwidth and first-load UX are real costs for your users. The claimed sub-10ms time-to-first-token is the author's own figure — verify it against your own hardware before designing a real-time autocomplete flow around it.

Discussion angle

Run the suite on every loaded model on your own laptop and compare the sustained tok/s and accuracy pass rates against the page's claims — then decide whether local SLM triage is good enough to replace a cloud classify call in one specific step of your agent or app pipeline.

Top