AI Weekly Malaysia

Back to items Summaries

Ember-1

ID
29188
Status
summarized
Published
28 Sep 2026, 1:31 AM
Fetched
28 Sep 2026, 6:07 AM
Provider
Hacker News
Category
dev-community
Original URL
https://fireworks.ai/blog/ember-1
Source URL
https://hnrss.org/best

Summary

Score
6.0
Created
28 Sep 2026, 6:07 AM
Tags
Audience
developersvibe_codersai_agent_usersai_ml_learners

What happened

Fireworks Research launched Ember-1, a specialized model built on Kimi K3 that it says matches K3's quality with 40% fewer tokens. The stated driver: K3's long reasoning traces (sometimes over 90% of generated tokens) made automated coding expensive at scale, and simply lowering reasoning effort gave up too much quality, so Fireworks retrained instead — running 50+ training experiments and 200+ evaluations on Fireworks Serverless Training. It was validated on public benchmarks, an internal 'Specialized Intelligence Index', Fireworks' own coding/agent workloads, and live customer A/B tests; it's the first in a planned series. The Hacker News thread drew 236 points and 132 comments.

Why it matters

If you run coding agents or long agent loops on Kimi K3 and pay per token, a claimed 40% token reduction is a direct cost lever — but every number here is Fireworks' own, with no independent benchmark figures in the post and no pricing shown, so treat it as a hypothesis to A/B on your own traffic before migrating. The more transferable detail is that low reasoning-effort settings were not an acceptable substitute for a retrained model, which argues against assuming you can just dial effort down to cut agent costs.

Discussion angle

Fireworks says lowering K3's reasoning effort lost too much quality, so they trained a new model instead — if that's true, 'just use low effort mode' is not a real cost strategy for agent workloads. How would you A/B a 40%-fewer-tokens claim on your own traffic, and what would you need to see (price, latency, failure rate) before switching a production coding agent off K3?

Top