AI Weekly Malaysia

Back to items Summaries

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

ID
31668
Status
summarized
Published
04 Oct 2026, 8:51 PM
Fetched
05 Oct 2026, 1:27 AM
Provider
Hacker News
Category
dev-community
Original URL
https://github.com/Niko1221/Strata
Source URL
https://hnrss.org/best

Summary

Score
7.0
Created
05 Oct 2026, 1:27 AM
Tags
Audience
developersai_ml_learnersai_agent_usersvibe_coders

What happened

Strata is an open-source, one-click inference engine (GitHub Niko1221/Strata, 10.2k stars, 902 forks, 845 commits) that runs the 125B-parameter Qwen3.8-Flash-Next on consumer GPUs with 12GB+ VRAM on Windows or Linux, exposing an OpenAI/Anthropic-compatible API on localhost with optional image input. Its own benchmark table shows Q2_0 hitting 94 tok/s generation and 2,650 tok/s prompt processing on an RTX 5070 12GB / Ryzen 5 7600 / 64GB RAM, and 60 / 1,160 tok/s on an RX 9070 XT 16GB / Ryzen 9 3900X / 47GB RAM, with quality dropping down the quantization ladder (IQ3_S: 53 / 1,620 tok/s). The Hacker News thread drew 265 points and 134 comments.

Why it matters

If you pay per-token for coding agents or chat, a localhost OpenAI-compatible endpoint on a 12GB card is worth a test — but the submission title claims '100T/s' and an RTX 4090, while the repo's own numbers top out at 94 tok/s on an RTX 5070, so treat the headline as unverified. Everything here is 2-bit-class quantization (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S), so benchmark your actual coding tasks against a hosted model before pointing a production agent at it; the speed cost of stepping up to IQ3_S is roughly 40 tok/s, which is the real tradeoff to decide on.

Discussion angle

Run the same coding task through Strata's localhost endpoint at Q2_0 and IQ3_S, then against whatever hosted model you use — does the quality gap at 2-bit quantization justify the 94 tok/s local speed, and does the advertised 128K context actually hold on a 12GB card without spilling to system RAM?

Top