October 7, 2026Open SourceAgentsCoding

Mistral Large 4: a 1T Open-Weight Model That Will Do the Cyber Work Closed Models Refuse

Mistral shipped a trillion-parameter model on Tuesday and Hacker News gave it 1,449 points before dinner. Mistral Large 4, nicknamed Le Chonk, is a natively multimodal mixture-of-experts model with 1T total parameters and 49B active. It is a public preview on Mistral Studio today. The weights land at the end of the month, after red-teaming with security firms and state authorities who get a less moderated version.

The headline number is not the size. It is where Mistral chose to compete. On the Artificial Analysis Cyber Index the model ranks top five globally, and on the test that asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model. Mistral says plainly why: Claude Opus 5.5 and GPT-6 Astra score near zero on that same test because they refuse to do it. The pitch to defenders is that losing a capability mid-incident is itself a security risk, and open weights under your own policy fix that. It is the sharpest open-versus-closed argument any lab has put in a launch post this year.

On agent work the numbers are good for an open model and not frontier. DeepSWE v1.1 at 61.7%, SWE-Atlas-QnA at 59.4%, Terminal-Bench 4 at 28.3%, a Coding Agent Index of 49.8% that puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. On AutomationBench, 657 business workflows across Gmail, Sheets, Slack and Salesforce, it scores 59.9%. In a blind Surge AI human eval of coding quality it came second of five at 3.74, behind only Claude Opus 5 at 4.22. Prompt-injection resistance is 93.3% on Lakera's B3 benchmark, which matters more for an agent model than most of the other rows.

The compute story is the European sovereignty angle. The model was trained from scratch on 3,800 Grace Blackwell GPUs in Mistral's own data centers, which VP of Science Pierre Stock told TechCrunch is two to three times fewer than Chinese competitors use. The training data spans 160+ languages. Samsung led a Series D last month at a 21 billion euro valuation, and ASML led the C, which explains why chip design and engineering drawings get their own benchmark section.

What to watch is the three-week gap. A 1T model that is preview-only is a benchmark claim. A 1T model with downloadable weights and an 82% exploit-reproduction score is a different kind of object, and the "trusted partners and governments" framing suggests Mistral knows it.

Link: mistral.ai/news/mistral-large-4/
← Previous
Ops Log: 2026-10-06
Next β†’
PolicyLM-1.7B: Decision Models Get Their First Real Job, Moderating Humans
← Back to all articles

Comments

Loading...
>_