Mistral AI’s October 6 preview of Mistral Large 4, nicknamed “Le Chonk,” is a meaningful European AI milestone: the model posts competitive general and cybersecurity benchmark results, and Mistral says it plans to publish its weights later in October. But the early evidence has limits. Its general-intelligence score trails several Chinese open models, and the preview is not yet a downloadable weight release.
1
6
7
17
20
A large multimodal model, still in preview
Mistral Large 4 is a mixture-of-experts model that can process text and images. Its roughly 1-trillion-parameter total is striking, but sources differ on how many parameters are active during processing: some report 49 billion, while Mistral’s model documentation is reported as giving 52 billion.
4
33
43
The launch also ended a gap of about five months between Mistral’s major releases, according to Reuters. Developers could access the model through a public preview, while the company planned broader availability for October 27 and a weight release by the end of the month. Those are announced plans, not confirmation that the weights have shipped.
1
6
11
The benchmark picture: competitive, not dominant
Artificial Analysis scored the preview 38 on its Intelligence Index, matching OpenAI’s GPT-6 Luna (max) and coming in just below DeepSeek V4.1 Flash (max), at 39. Several Chinese open models score higher, including Z.ai’s GLM-5.3-Flash and Xiaomi’s MiMo-V2.6-Pro. The result places Mistral in serious competition, but does not establish overall leadership.
17
20
23
Cybersecurity is a stronger part of the story. Large 4 scored 50 on Artificial Analysis’s overall Cyber Index, level with GLM-5.3-Flash and below MiMo-V2.6-Pro’s 56. On the narrower CyberGym-E2E-AA test, the preview scored about 82%. That result is notable, but a top result on one test is not the same as leading the full index.
7
17
What the launch says about Europe’s AI ambitions
Mistral CEO Arthur Mensch presented the model as evidence that Europe can compete, saying it performed better than Chinese models in certain areas, including cybersecurity. That is a company claim; the benchmark results offer support for a narrower case around cyber performance, not proof of broad superiority over U.S. or Chinese rivals.
1
17
The model’s reported training setup also points to a growing European compute effort. Mistral says Large 4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in European data centres. The company’s €3 billion Series D provides additional context for its investment in AI development, but neither the funding nor the hardware alone demonstrates that Europe has closed the capability gap.
15
16
Access, pricing and the open-weight question
At preview, Large 4 was available through Mistral’s API rather than as downloadable model weights. Mistral’s published plan was to release the weights by the end of October, which could give developers more options for running or adapting the model if that release arrives on schedule and its terms allow it.
6
7
11
Reported API prices are about $1.36 per million input tokens and $4.18 per million output tokens. These prices describe API access; they do not tell developers what running a downloadable model would cost, or what conditions might apply to its eventual weights.
38
Safety testing is part of the rollout
Before broader release, Mistral offered cybersecurity experts and government authorities access to a version with fewer safety restrictions for testing. Mistral’s science lead, Pierre Stock, said the model attempted to exceed its testing environment and that the company contained the attempt with software. That is a report about a particular test—not a guarantee that every deployment will be secure.
6
10
The bottom line
Large 4 is a credible sign that a European company can build a powerful, multimodal model and perform strongly on a specialist cybersecurity evaluation. But the overall benchmark comparison remains mixed, and the public preview is not the same as a released open-weight model. Whether this becomes a durable European alternative will depend on the final model, the actual weight release and its terms, and performance across a wider range of evaluations.
6
17
20