AI Weekly Malaysia

Back to items Summaries

Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs

ID
22774
Status
summarized
Published
09 Sep 2026, 4:07 AM
Fetched
09 Sep 2026, 9:52 PM
Provider
Hacker News
Category
dev-community
Original URL
https://github.com/argonautlabsai/deltafin
Source URL
https://hnrss.org/best

Summary

Score
7.0
Created
09 Sep 2026, 11:02 PM
Tags
Audience
developersai_ml_learnersvibe_coders

What happened

Argonaut Labs forked gavamedia's deltafin engine to run Kimi K3 (2.8T-parameter MoE, 1.45 TB expert weights) on a single M5 Max MacBook Pro with 128 GB RAM, streaming experts from four SSDs. It achieves ~1 token/s steady decode, but a 512-token prompt takes ~6.3 minutes to first token due to prefill re-reading each layer's experts 8x. Drive scaling shows diminishing returns: one drive gives 52% of four-drive speed, two 73%, three 90%.

Why it matters

If you're experimenting with running frontier-scale MoE models locally, this shows the bottleneck is expert weight I/O, not compute—and that adding SSDs has steep diminishing returns because the slowest of each layer's 16 reads sets the pace. The prefill bottleneck (8x re-reads) is identified but not yet fixed, so don't expect interactive latency yet.

Discussion angle

What does the drive-scaling curve (52/73/90/100%) tell us about architecting storage for local MoE inference—is it worth building a 4-SSD rig, or does the slowest-read bottleneck make it uneconomical versus just renting cloud GPUs?

Top