DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- ID
- 22827
- Status
- summarized
- Published
- 09 Sep 2026, 7:19 PM
- Fetched
- 11 Sep 2026, 6:13 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://news.ycombinator.com/item?id=49624603
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.5
- Created
- 11 Sep 2026, 6:14 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_users
What happened
DeepSeek is releasing V4.1 Flash around September 10, 2026 (Beijing Time), claiming it surpasses V4 Pro on performance, cost, speed, and task completion. Critically, all API requests to the Pro model will be automatically rerouted to V4.1 Flash at Flash pricing until V4.1 Pro ships—meaning existing Pro workflows get a forced model swap. Off-peak pricing is $0.003 for input cache hits, $0.15 for cache misses, and $0.6 for output, with peak hours at double those rates.
Why it matters
If you have production workflows validated on V4 Pro, DeepSeek is forcing a model swap you didn't ask for—test V4.1 Flash now before September 10 or migrate to an alternative open-weights provider like Together.ai or OpenRouter to keep running V4 Pro. The peak/off-peak pricing split also means you should batch non-urgent inference jobs to off-peak hours to halve your costs.
Discussion angle
The forced rerouting from Pro to Flash is a cautionary tale for anyone building on hosted LLM APIs—what's your fallback plan when a vendor swaps your model out from under you, and does open-weights availability change your provider selection criteria?