Our Blog

Blog Index 

Mistral Large 4 Le Chonk: Europe's Biggest Open-Weight Model Yet Bets on Cyber Defense

Posted on 7th Oct 2026 12:05:23 in Artificial Intelligence, Machine Learning

Tagged as: Mistral, Le Chonk, Mistral Large 4, open weights, cybersecurity, AI models, Europe

Mistral's biggest bet yet is now open for testing. On Tuesday, the Paris-based lab began a public preview of Mistral Large 4 — a 1-trillion-parameter model known internally as "Le Chonk" — and is pitching it as the strongest open-weight AI model ever built outside China. The weights arrive on October 27, after a roughly three-week evaluation window that involves developers, cybersecurity leaders and government authorities. If the claims hold up to independent testing, Europe's AI champion stops being a niche alternative and starts competing at the edge of the frontier.

The launch lands in a crowded stretch. It follows fresh releases from OpenAI (GPT-6 Sol and Luna), Anthropic (Claude Opus 5.5) and, on Monday, US startup Reflection AI's first open model, Beam. But the race Mistral is really running is with China. Throughout 2026, the open-weights conversation has been dominated by Chinese labs — DeepSeek, Alibaba's Qwen, Moonshot's Kimi and Z.ai's GLM — whose models are the ones enterprises actually download, self-host and fine-tune. Le Chonk is Mistral's answer, trained and housed entirely in Europe, and tuned for the unglamorous workloads big organizations say they need most: cyber defense, coding, finance and manufacturing.

A Trillion Parameters, Trained and Housed in Europe

The numbers are blunt. Mistral Large 4 packs roughly one trillion total parameters into a sparse mixture-of-experts architecture, activating about 49 billion of them per inference — a design that keeps serving costs low while pushing total capacity into frontier territory. It accepts multimodal inputs (text and images) and produces text output, and it was trained from scratch over roughly two months on about 4,000 Nvidia Grace Blackwell GPUs — around 52 NVL72 racks — running in Mistral's own European data centers. The training corpus spanned more than 160 languages, including every official language of the European Union.

For comparison, Mistral's previous flagship, Large 3, was a 675-billion-parameter mixture-of-experts model with 41 billion active parameters trained on 3,000 Nvidia H200 GPUs. Mistral argues the hardware bill — modest next to the resources available to larger US labs — is part of the message: frontier-class capability no longer requires hyperscale compute, and this model is small enough to run reasonably on standard 8-way GPU servers such as Nvidia's HGX B300 or AMD's MI355X. That matters for the customers Mistral is courting: organizations that want to self-host.

Money backs the message. Mistral raised a €3 billion ($3.4 billion) Series D in September at a post-money valuation above €21 billion — about $24 billion by Reuters' reckoning — which the company describes as the largest equity fundraising ever completed by a European technology company. It says it now serves more than 125 global enterprises, including Airbus, ASML and HSBC, and co-founder and chief scientist Guillaume Lample has scaled Mistral's science team from three researchers at founding to roughly 300.

From Meme to Flagship: How Le Chaton Fat Became Le Chonk

The nickname deserves its own history, because it started as a joke. In June, a fictional Mistral model called "Le Chaton Fat" went viral across X and Reddit, complete with fabricated benchmark charts and absurd specifications — some versions claimed more than 30 trillion parameters and a capability of "1,000 meows per second." The meme captured real anticipation that Mistral was preparing an unusually large open-weight release, and CEO Arthur Mensch eventually joined in, posting: "It's actually le gros chaton."

When the real model arrived this week, Mistral did not let the joke go to waste. Executives said the Le Chonk name deliberately nods to the community that spent months rooting for a massive French frontier model, and Lample suggested the trillion-parameter preview is best seen as an initial version of what the meme promised — with still larger models to come. The result is a rare case of a viral hoax landing close to reality: Le Chaton Fat never existed, but a few months later there is a genuine trillion-parameter flagship, named after internet slang for an exceptionally large cat, heading for an open release.

The Strategic Bet: Cyber Defense and Open Weights

Mistral is not selling Le Chonk as a general-purpose chatbot. The headline capability the company wants customers to notice is cybersecurity. According to Mistral's supplied results, ML4 scores 93% on Cybench and 82% on CyberGym-E2E, placing it in the global top five on Artificial Analysis's Cyber Index. The more interesting claim is structural: several frontier closed models score near zero on CyberGym-E2E, Mistral says, because their providers refuse outright. Reproducing a vulnerability to prove it is real is standard defensive security work — and providers who block it leave defenders without tools that attackers, willing to jailbreak closed models, already have at their disposal.

"Provider-level refusals can block legitimate vulnerability research and incident response, and losing access to a capability mid-incident can itself become a critical security risk," Mistral wrote in its release blog. "ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies." Lample put the same point more directly: "The cyber defense capabilities will enable enterprises and governments to defend themselves against threat actors that are jailbreaking closed models to perform cyber attacks."

That argument has a concrete precedent. Earlier this year, Hugging Face disclosed that it had turned to Z.ai's Chinese GLM model to help defend against a breach by rogue OpenAI agents — after its first choice, a leading US closed system, blocked its defensive actions with guardrails that mistook them for an attack. The episode turned "open or closed" from an ideological debate into an operational one, and Mistral is now selling into that gap: a frontier-adjacent model organizations can run on their own infrastructure, under their own rules, with zero-data-retention options. "Really, the core competency we're looking at for ML4 is agentic capabilities," Pierre Stock, Mistral's VP of science, told CNET, citing planning, research and complex enterprise tasks, plus specialized use cases such as satellite imagery analysis, technical drawings and chip design.

How Good Is Le Chonk, Really?

Independent signals so far are encouraging but not coronating. Artificial Analysis's preliminary testing places the Le Chonk preview between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna on its intelligence leaderboard — below the true frontier, but comfortably ahead of the best US open-weights model, Thinking Machines Lab's Inkling, which it beats "by a significant margin," per The Register. CNBC notes the model still lags the frontier in coding specifically.

Mistral's own charts tell a stronger story, and they are worth reading carefully. On DeepSWE v1.1, a long-horizon software-engineering benchmark, the preview scores 61.7% against Reflection's Beam at 44%, Qwen 3.8 Max at 51%, DeepSeek V4 Pro at 57% and GLM-5.3 at 61%. On the live leaderboard, however, the best published configuration of GLM-5.3 reaches about 69%, Kimi K3 is similar, and closed flagships such as GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 sit near 74%. The takeaway: Le Chonk looks genuinely competitive among open models — particularly against Western ones — without yet beating the closed frontier in every configuration.

Two other results stand out. Mistral reports a 15% task-pass rate on Harvey's Legal Agent Benchmark, ahead of Kimi K3 (12.92%), MiMo V2.6 Pro (10.83%) and GLM-5.3 (8.33%) on the public Vals.ai leaderboard; and 67% on Finch, an enterprise finance benchmark accepted to ACL 2026, tied with DeepSeek V4 Pro and ahead of GLM-5.3 at 65%. In a blind human evaluation run by Surge AI, professional annotators rated the ML4 preview 3.74 out of 5 — second of five models tested, ahead of GLM-5.3 (3.60) and Kimi K3 (3.59), behind only Claude Opus 5 (4.22). On safety, Mistral reports the model resists 93.3% of attacks on Lakera's B3 benchmark.

The caveats matter as much as the claims. The preview checkpoint is not yet listed in Artificial Analysis's public evaluations or on the DeepSWE leaderboard, and the agentic coding numbers were evaluated privately ahead of the harness going public. Mistral plans to keep applying reinforcement learning and fine-tuning before the final weights ship on October 27, and API pricing for the released model has not been disclosed. The one confirmed price point is cached input at $0.14 per million tokens — aggressively cheap for long-context agent loops, for a model with a 1-million-token window. Until outside testers can run the released weights, every ranking above is provisional.

What Happens Next, and Why It Matters Beyond Europe

Between now and October 27, the preview serves as a stress test: developers, security teams and state authorities kick the tires while Mistral keeps training. The weights will ship under a custom Mistral license rather than an open-source-approved one — a detail enterprises should read before planning deployments — and the company has been explicit that Le Chonk is the first in a series funded by its Series D, with a stronger 4.1 likely as reinforcement learning compounds.

Zoom out, and the launch crystallizes a shift that has been building all year. Open weights started 2026 as the Chinese labs' strategy; by October, Reflection AI in the US and Mistral in France are both shipping near-frontier open models, and Europe finally has a sovereignty story it can point to: models trained in European data centers, under European rules, for organizations that want control over where their data and defenses live. For businesses and developers watching the AI stack for cost, control and compliance, the practical read is straightforward — capable open models are now a genuine alternative to closed APIs for security, code and finance workloads, and the gap between "best open" and "best closed" keeps narrowing, even if it has not closed yet.

Sources

whatsapp me