Mistral Large 4 Trails China’s Open-Weight Leaders on Artificial Analysis
Set Trending Topics as a preferred source on Google.
The first independent test is in, and it only half backs up Mistral’s claims: On Artificial Analysis, one of the most closely watched benchmarking services for A.I. models, Mistral Large 4 scores 38 points on the Intelligence Index. That makes it the strongest open model outside China, but seven Chinese models still rank ahead of it among open-weight models.
Mistral just unveiled the model and, in its launch blog post, pitched it as competitive with the strongest open models worldwide. The Artificial Analysis Intelligence Index combines the results of 10 evaluations covering areas such as coding, agents, knowledge work and science into a single score, and is considered one of the industry’s most important independent yardsticks.
Eighth Place, Behind Seven Chinese Models
With 38.4 points, Large 4 is a huge leap over its predecessors: Mistral Medium 3.5 scored 14 points and Large 3 scored 9. Because the model weights have not been released yet, Artificial Analysis currently lists the preview as a proprietary model. Among models in its price range, it ranks 64th out of 225, well above the median of 26 points.
The more telling comparison is with the open models in the Artificial Analysis ranking, the group Large 4 is set to join once its weights are released at the end of October. There, it would come in eighth. All seven models ahead of it are Chinese, led by Xiaomi’s MiMo-V2.6-Pro with 46.3 points, Z.ai’s GLM-5.3 with 44.8 and Moonshot AI’s Kimi K3 with 43.6. Until now, the strongest open model outside China was Motif 3 from South Korea, with 33.6 points. Mistral’s claim that Large 4 clearly outperforms every open model from the United States and Europe therefore holds up. Large 4 also beats GLM-5.2 (33.7 points) and DeepSeek V4 Pro (36.0 points), the models Mistral pointed to at launch. And the model may still improve before the weights are released at the end of October, since Mistral is continuing to train the preview.
The gap to the leading closed models remains wide: Anthropic’s Claude Opus 5.5 scores 57.6 points in its strongest setting, OpenAI’s GPT-6 Astra 52.7 and Google’s Gemini 4 Argon 52.6. That leaves Large 4 more than 19 points behind the world’s best, reaching about two-thirds of Claude Opus 5.5’s score. Even the best open model, MiMo-V2.6-Pro, is almost 8 points ahead.
The Open Models Compared
The top open-weight models on the Artificial Analysis Intelligence Index, plus Mistral Large 4 and the cost per task on the index:
| Model | Developer | Index | Cost per task (USD) | Cost per task (EUR) |
|---|---|---|---|---|
| MiMo-V2.6-Pro | Xiaomi | 46.3 | $0.13 | €0.12 |
| GLM-5.3 (Max) | Z.ai | 44.8 | $2.01 | €1.78 |
| Kimi K3 (Max) | Moonshot AI | 43.6 | $2.00 | €1.77 |
| GLM 5.3 Flash | Z.ai | 41.8 | $0.25 | €0.22 |
| Qwen3.8 2.4T | Alibaba | 39.9 | $2.16 | €1.92 |
| Qwen3.8-Flash-Next | Alibaba | 39.8 | $0.37 | €0.33 |
| DeepSeek V4.1 Flash (Max) | DeepSeek | 39.5 | $0.27 | €0.24 |
| Mistral Large 4 (Preview) | Mistral | 38.4 | $1.13 | €1.00 |
| MiMo-V2.6-Flash | Xiaomi | 37.9 | $0.06 | €0.05 |
| DeepSeek V4 Pro 0813 (Max) | DeepSeek | 36.0 | $0.67 | €0.59 |
Cheaper Than GLM and Kimi, Pricier Than DeepSeek
On cost per task, the picture is mixed. At $1.13 (about €1), Large 4 is a good 40 percent cheaper than its big rivals GLM-5.3, Kimi K3 and Qwen3.8, which each cost around $2 per task but also score several points higher. Against China’s more efficient models, however, Large 4 looks dated: DeepSeek V4.1 Flash scores slightly higher at $0.27 per task, and MiMo-V2.6-Pro delivers eight more points for $0.13.

One reason for the cost is how talkative the model is. Across the full Intelligence Index, Large 4 generated about 200 million output tokens, compared with a median of 81 million for comparable models. Artificial Analysis therefore rates it as “very verbose.” The complete test run cost about $1,600 (roughly €1,420). On speed, though, Large 4 does well: At 116 tokens per second, it beats the median of 87, and the first response arrives after 1.46 seconds.
What the Preview Can’t Do (Yet)
Anyone who wants to use Large 4 now has to accept several limitations:
- Mistral only: The preview is available exclusively through Mistral’s own API; according to Artificial Analysis, there is currently just one provider. The API costs $1.36 (about €1.21) per million input tokens and $4.18 (about €3.71) per million output tokens, with a discount of about 90 percent on cached inputs.
- No weights: The model can only be downloaded and run on private servers starting at the end of October. With one trillion parameters, 49 billion of them active, it is a job for data centers, not laptops, anyway.
- Throttled cyber capabilities: The public version is more heavily moderated. Until the release, selected security companies, partners and government agencies are testing a variant with fewer restrictions and expanded cyber features. Mistral also says it monitors traffic through the API.
- Text output only: Large 4 understands images and text but responds only in text. Its context window covers about 524,000 tokens.
- Scores in flux: According to Mistral, the reinforcement learning run is still underway. The Artificial Analysis scores are therefore a snapshot and are likely to change by the final version.
An Open Question: The License
It is still unclear under what terms the weights will be released. Mistral’s launch blog post talks at length about open weights and giving customers control, but does not name a license. According to VentureBeat, Large 4 will come under a custom Mistral license, with no details so far. That is more than a formality: Mistral has used very different licenses in the past. Large 3 was released under the permissive Apache 2.0 license, while Medium 3.5 came under a modified MIT license that excludes companies with more than $20 million in monthly revenue from free commercial use. Older versions of Mistral Large, by contrast, were released under the Mistral Research License, which ruled out commercial use without a separate agreement.
Terms vary among the Chinese competitors as well. DeepSeek and Xiaomi release their top models under the MIT license, while GLM-5.3, Kimi K3 and the large Qwen3.8 come with their own licenses, which Artificial Analysis classifies as restricted for commercial use. For companies that want to run Large 4 in their own data centers, the license will determine how open the model really is in practice.

