» That one asked for a bit of basic preparation. Trained and served on our own compute infrastructure, and RL shows no signs of saturation. » On X, Arthur Mensh plays it modest to present Mistral Large 4 (ML4), aka The Chonk.
With Mistral Large 4 (ML4), Mistral AI does not claim to dethrone the most advanced closed models right away. A general‑purpose multimodal model built on a granular Mixture‑of‑Experts (MoE) architecture, its documentation states 1.05 trillion parameters in total, of which 52 billion are active per token, as well as a 1.6‑billion‑parameter vision encoder.
The launch brief, for its part, mentions 49 billion active parameters.
Read also: Mistral and Cloudera seal a partnership for a sovereign AI
This MoE architecture enables only a fraction of the model to be mobilized at each generation step. The goal is to combine very high potential capacity with a more controlled inference cost than a dense model of comparable size.
A very wide context window
The primary usage promise of ML4 lies in its context window announced at 1 million tokens. It should allow handling very long documents, software repositories, regulatory corpora, or enterprise knowledge bases without systematically splitting them. And it supports more than 160 languages, including all official languages of the European Union.
The model also exposes usage‑oriented features such as function calls, structured outputs, document-based question‑answering, as well as batch processing, agents, and conversations. In other words, Mistral positions ML4 as a building block for enterprise applications and workflows.
Strong benchmarks, without domination
On Artificial Analysis’s Intelligence Index, ML4 scores 38 points.
C’est un bond important par rapport à Mistral Large 3, crédité de 9 points, et cela place le modèle au niveau de GPT-6 Luna sur cet indicateur composite.
| Indicator | Mistral Large 4 | Interpretation |
|---|---|---|
| Intelligence Index | 38 | At the level of GPT-6 Luna, behind top closed models. |
| Cyber Index | 50 | Equal to GLM-5.3-Flash, behind MiMo-V2.6-Pro. |
| DeepSWE v1.1 | 61.7% | Solid result reported in agentic coding. |
| Terminal-Bench 4.0 | 28.3% | Relatively weak score on terminal‑agent tasks. |
| CyberGym-E2E-AA | 82% | Ahead of GPT-6 Luna (78%) and MiMo-V2.6-Pro (79%). |
| Generation speed | 116.1 tokens/s | Above the median measured by Artificial Analysis. |
| First token latency | 1.46 s | Lower than the median of the compared reasoning models. |
But performance is not uniform: the 28.3% score on Terminal-Bench 4.0 shows ML4 lags behind in some demanding automation scenarios.
The figures released by Mistral should also be read with caution, as the model is still in pre‑release and its weights are not publicly available yet.
Against closed models
ML4 sits among the top tier of non‑American and non‑Chinese models but far from the market’s summit.
Read also: Mistral and HUMAIN bet on Arabic‑speaking AI models
Claude Opus 5.5 scores 58 points on the Intelligence Index, versus 53 for GPT-6 Astra and Gemini 4 Argon, while ML4 stops at 38.
| Model | Type | Intelligence Index | Cost per task | Context |
|---|---|---|---|---|
| Claude Opus 5.5 | Closed | 58 | $5.98 | 1 M |
| GPT-6 Astra | Closed | 53 | $3.26 | 1 M |
| Gemini 4 Argon | Closed | 53 | $1.99 | 1 M |
| GPT-6 Luna | Closed | 38 | $0.07 | 1 M |
| Mistral Large 4 | Open weights announced | 38 | $1.13 | 524 k measured by Artificial Analysis |
ML4 thus matches GPT-6 Luna on overall intelligence, but costs about 16 times more per task according to Artificial Analysis measurements. Its edge is therefore not API usage price but the promise of open weights and the control they offer to businesses.
The decisive advantage: open weights
Mistral announces that ML4’s weights will be published by the end of October. This decision differentiates the model from OpenAI, Anthropic, and Google offerings, which remain largely accessible via API.
Open weights allow an organization to deploy ML4 in its own private cloud, its VPC, or on its servers. It can thus keep data within its perimeter, set its own security policies, and tailor the model to its business needs through fine‑tuning.
Mistral presents ML4 as a way for its customers to retain control over intelligence, data, compute, and operations.
A clear business positioning
ML4 targets three main segments:
- Businesses that want to leverage a high‑performing multimodal model without sending their data to a third‑party API.
- Organizations with needs in cybersecurity, code analysis, or workflow automation, areas where ML4 delivers competitive results.
- Public and industrial actors sensitive to European digital sovereignty, for whom hosting in Europe and control over the weights are decisive factors.
The launch API pricing is $0.68 per million input tokens, $0.07 per million input tokens in cache, and $2.09 per million output tokens, after a 50% discount during the first two weeks.
Read also: Mistral AI changes the substance and form for content moderation
What ML4 really brings
ML4 brings three things to the LLM landscape:
- A European, high‑capacity multimodal model with a one‑million‑token context window and agent‑oriented capabilities.
- Credible performance in cybersecurity and agentic coding, not yet matching the top closed models on every indicator.
- An open‑weights sovereignty alternative, enabling enterprises to deploy, customize, and control the model within their own infrastructure.