Mistral Large 4 “Le Chonk”: 1.05 trillion parameters in public preview
Mistral opened the public preview of Mistral Large 4 on 6 October, its first big frontier model in months. It comes with open weights, leans hard into cybersecurity, and is aimed at Chinese models before American ones.

On 6 October, Mistral opened a public preview of Mistral Large 4, its new frontier model, nicknamed “Le Chonk”. For months the French company had said very little about big models. Now it’s back with a heavyweight, and the nickname is honest about what the scales say. The details below come from the presentation relayed by Next. Most of the performance figures are Mistral’s own, so you’ll want to check them against the final model.
A 1,050-billion-parameter MoE
On paper, the model has 1,050 billion parameters, with 49 billion active at any one time. It uses a Mixture of Experts architecture, which is now standard: a huge committee of experts where only a few speak up on each question. Think of a meeting where most people stay quiet and someone actually gets to the point. Mistral announces a context window of one million tokens. It says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell chips, in its own data centres in Europe.
Take the word “preview” literally: the model is still training. Until the final release, scheduled for 26 October, selected partners get a private version. It gives them lighter moderation, more advanced cybersecurity features and the preview weights. Mistral promises public weights and more detailed information for the same date, 26 October.
Flattering benchmarks, a bill worth watching
Mistral claims 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. That adds up to a combined index of 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. Artificial Analysis gets the same scores and ranks the model in the upper-middle of the pack, for both intelligence and speed. It also lists it among the very wordy models. The pricing looks reasonable: $1.36 per million input tokens and $4.18 per million output tokens. But when a model loves the sound of its own voice, output tokens are the ones that count, and your invoice will notice before you do.
There is one inconsistency. Artificial Analysis lists a context window of 524,000 tokens, about half the announced figure. Next suggests that number may come from an earlier preview.
Cybersecurity as the sales pitch
This is where Mistral pushes hardest. The company calls Large 4 “one of the most powerful AI models in the world for cybersecurity”. It reportedly ranks in the Top 5 of Artificial Analysis’s Cyber Index with 82%, and scores 93% on Cybench. Mistral is blunt about its American rivals: Claude Opus 5.5 and GPT-6 Astra score zero on the Cyber Index because they refuse to perform the tasks. According to Mistral, attackers already get around these closed models with jailbreaks, and “defenders need systems that can match these capabilities without being constrained by the same refusals”.
Next points to the Hugging Face case. When its services came under attack, the company complained that closed models refused to help it analyse the first traces. So it turned to a Chinese model, GLM 5.2. Mistral presents all this as a question of defence. The brochure says nothing about what a more open, less restricted model also gives attackers.
Up against the Americans, and above all the Chinese
Mistral calls its model “by far the best in Europe and the United States” among open-weight models. That choice of regions says a lot: for open weights, the models to beat are still Chinese, and Mistral’s comparisons target exactly DeepSeek and Qwen. For office work, Mistral announces 59.9% on AutomationBench. The model reportedly beats GPT-6 Astra on financial and legal tasks evaluated by vals.ai, and leads the open models on Harvey’s legal test. In visual grounding, it scores 42% on Dense 200 against 41% for GPT-6 Astra. One point apart is a photo finish.
Next wonders whether Le Chonk will soon replace GLM 5.3 in Vibe Code. That would be the strongest vote of confidence: Mistral eating its own cooking by running its own coding tool on its own model.
Sources (1)
Written with AI assistance from the sources cited above, then reviewed and approved before publication by Sébastien Soulier.


