AI Weekly Malaysia

Back to items Summaries

The AI Race Just Got Awkward

ID
30467
Status
summarized
Published
30 Sep 2026, 11:50 PM
Fetched
01 Oct 2026, 3:20 AM
Provider
Hacker News
Category
dev-community
Original URL
https://insufferable.dev/posts/the-ai-race-just-got-awkward/
Source URL
https://hnrss.org/best

Summary

Score
6.0
Created
01 Oct 2026, 3:21 AM
Tags
Audience
developersai_ml_learnersai_agent_usersstartup_founders

What happened

A blog post on insufferable.dev argues the competitive dynamic between Western and Chinese AI labs has flipped: instead of Western labs accusing Chinese labs of distilling their models, Western labs are now quietly adopting Chinese inference optimizations. It cites DeepSeek's KV cache work — MLA at roughly 15x compression, then Compressed Sparse Attention and Heavily Compressed Attention, and DeepSeek-V4.1-Flash with CSA2, cross-layer cache reuse and FP4 caching bringing the global KV cache to 890 bytes per token, roughly 437x below DeepSeek-V1 — and claims Claude Opus 5.5 and GPT-6.1 Sol shipped with these techniques, with Opus 5.5 cutting cache-read pricing 60% versus Opus 5. The excerpt is truncated mid-sentence, and the pricing claims and model-release details are asserted by the author without cited primary sources.

Why it matters

If the cache-read price cuts described here are real, the cost of running long-context coding and agent sessions shifts from output tokens toward a much cheaper cache-read line item, which changes how you'd budget and architect retrieval-heavy agents. But the article gives no links to DeepSeek's papers or to Anthropic/OpenAI pricing pages, so before repricing anything, verify the 890 bytes-per-token figure and the claimed 60% Opus cache-read reduction against the vendors' own docs — the HN thread (349 points, 368 comments) is a better starting point than the post itself.

Discussion angle

Pull your own LLM bill and split it into cache-read versus output tokens for a long-context agent workflow — does a 60% cache-read cut actually move your total, or is output still the dominant cost?

Top