Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
- ID
- 31668
- Status
- summarized
- Published
- 04 Oct 2026, 8:51 PM
- Fetched
- 05 Oct 2026, 1:27 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://github.com/Niko1221/Strata
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 05 Oct 2026, 1:27 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coders
What happened
Strata is an open-source, one-click inference engine (GitHub Niko1221/Strata, 10.2k stars, 902 forks, 845 commits) that runs the 125B-parameter Qwen3.8-Flash-Next on consumer GPUs with 12GB+ VRAM on Windows or Linux, exposing an OpenAI/Anthropic-compatible API on localhost with optional image input. Its own benchmark table shows Q2_0 hitting 94 tok/s generation and 2,650 tok/s prompt processing on an RTX 5070 12GB / Ryzen 5 7600 / 64GB RAM, and 60 / 1,160 tok/s on an RX 9070 XT 16GB / Ryzen 9 3900X / 47GB RAM, with quality dropping down the quantization ladder (IQ3_S: 53 / 1,620 tok/s). The Hacker News thread drew 265 points and 134 comments.
Why it matters
If you pay per-token for coding agents or chat, a localhost OpenAI-compatible endpoint on a 12GB card is worth a test — but the submission title claims '100T/s' and an RTX 4090, while the repo's own numbers top out at 94 tok/s on an RTX 5070, so treat the headline as unverified. Everything here is 2-bit-class quantization (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S), so benchmark your actual coding tasks against a hosted model before pointing a production agent at it; the speed cost of stepping up to IQ3_S is roughly 40 tok/s, which is the real tradeoff to decide on.
Discussion angle
Run the same coding task through Strata's localhost endpoint at Q2_0 and IQ3_S, then against whatever hosted model you use — does the quality gap at 2-bit quantization justify the 94 tok/s local speed, and does the advertised 128K context actually hold on a 12GB card without spilling to system RAM?