Mistral Large 4
- ID
- 32362
- Status
- summarized
- Published
- 06 Oct 2026, 9:15 PM
- Fetched
- 06 Oct 2026, 11:41 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://docs.mistral.ai/models/mistral-large-4-0
- Source URL
- https://hnrss.org/best
Summary
- Score
- 8.0
- Created
- 06 Oct 2026, 11:42 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_userssaas_startup_founders
What happened
Mistral published docs for Mistral Large 4, an open-weight multimodal model in Public Preview as of October 6, 2026, built on a granular Mixture-of-Experts architecture with 49B active parameters, 1.05T total parameters, and a 1.6B vision encoder. It lists a 1M-token context window and pricing of $1.36/$0.68 per M input tokens, $0.14/$0.07 per M cached input tokens, and $4.18/$2.09 per M output tokens (the lower figure in each pair appears to be the batch rate). The docs page covers structured outputs, function calling, document QnA, prefix, chat completions, batching, agents/conversations endpoints, and built-in tools, but shows no benchmark numbers, license terms, or weight download links. The Hacker News thread drew 528 points and 285 comments.
Why it matters
The decision this changes is your model-routing default: a 1M-context multimodal model at $0.68/M input and $2.09/M output is cheap enough to move long-document and multi-turn agent workloads off short-context models, and because weights are open, you can weigh self-hosting against API cost when data residency or per-token spend matters. Before switching, note what the page does not give you — no benchmarks, no license text, no weight links — so treat it as a pricing/spec claim to validate on your own eval set rather than a drop-in replacement.
Discussion angle
49B active out of 1.05T total is an aggressive MoE sparsity ratio — discuss what that means for serving cost and latency versus a dense model of similar quality, and whether the listed $0.68/$2.09 batch rates are realistic for the long-context agent loops you actually run.