On October 6, Mistral launched a public preview of its new model with an informal codename — le Chonk, a fat cat. The spec sheet is anything but casual: a 1-trillion-parameter mixture-of-experts (MoE) model with 49B active parameters per token and natively multimodal input. The company has committed to releasing the open weights by the end of the month (Mistral's official blog).
Trained from scratch on European soil
ML4 does not run on someone else's cloud: it was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and the preview API is served on that same infrastructure. Its training data covers more than 160 languages, including every official language of the European Union. The model will be available across multiple regions, including a European deployment that Mistral operates end-to-end, independently of other digital service providers and under European law — AI sovereignty is the core narrative of this release, and the first milestone after its €3 billion Series D.
Cybersecurity: where closed models refuse to answer
The boldest numbers are in cybersecurity. On the Artificial Analysis Cyber Index, an independent evaluation, ML4 ranks among the top five models globally and leads open-weight models developed outside China. On a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82% — per the official blog, the highest of any model. It also solves 93% of the 40 challenges in Cybench. The most striking contrast: Claude Opus 5.5 and GPT-6 Astra score near zero on the same test because their safety policies make them refuse the task. Mistral's argument is blunt — defending software often starts with proving a flaw is real, exactly the step where closed-model safety filters get in the way. During red-teaming, the company gave cybersecurity leaders, vetted partners, and state authorities access to the same model with reduced moderation and expanded cyber capabilities. On the safety side: 93.3% resistance to indirect prompt injections on Lakera's B3 benchmark, and a refusal rate on malicious cyber prompts higher than all open-source models.
Coding and agents: benchmarking against Chinese flagships
On coding: 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, and a combined Coding Agent Index of 49.8%, which Mistral says places it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind human evaluation by Surge AI, ML4 ranked second of five models at 3.74, behind Claude Opus 5 (4.22) and ahead of Kimi K3 (3.59) and GLM-5.3 (3.60). On the agentic side, it scores 59.9% on AutomationBench's 657 real business workflows, which the company says beats Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro; on visual grounding (Dense 200), it posts 42% against GPT-6-Astra's 41%. API pricing is $1.36 per million input tokens and $4.18 per million output tokens.
Post-training and what comes next
The post-training details are informative too. At a scale of 3k GPUs, a single RL training run produces roughly 33 billion tokens of rollouts per day, of which around 16 billion are trainable after filtering. Mistral says training rewards are still climbing with no signs of saturation. Third-party reports suggest the weights will go public around October 27 (TNW). The reality then: a 1T-total-parameter MoE, even with only 49B active, has a VRAM footprint that makes it an enterprise self-hosting proposition. For the first time, the open-weight narrative is shifting from "matching closed models" to "doing what closed models won't" — cybersecurity, sovereign deployment, controllable moderation. For enterprises, this is Europe's first trillion-scale open-weight option; for developers, wait for the weights to land, then let your GPU do the talking.