2026-10-06 · 24 min · mixture-of-experts · multimodal · benchmarks · security
Why read this
Notabletop 60%Checks Large 4's 'best from US or Europe' against Artificial Analysis records (true; eighth overall) and its charts against its prose; plus a memory planner.
- Checked against the source
- Widely used
- Concrete numbers to act on
LLM architectureAPI onlyPractitioner model
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 3 of 3: Reproduces the headline result, or shows from primary files it is wrong
- Can I run it?
- 1 of 3: API-only, gated or restrictive licence
- Will I understand it?
- 1 of 3: Partial mechanism
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 2 of 3: A widely used model, tool or lab release
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 59 of 100, ranked 249 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
For about a year, "Le Chonk" was a joke people made at Mistral: a fat cat standing in for the big model everyone assumed Mistral was not going to ship. This morning Mistral shipped it under that name. The post makes four claims in four lines: "1T parameters, natively multimodal. 49B active"; "the best open weights model from US or Europe on aggregated benchmarks"; "state-of-the-art on critical workloads, including cyber defense, manufacturing and finance"; and that it "surpasses closed frontier models on visual grounding". The API is live today. "Open weights release end of October."
I wanted to know two things. Is the "best from US or Europe" line true, and which aggregate is it talking about? And since there are no weights, how much of the rest can anyone check today?
The short version surprised me twice. The headline claim holds, and by a wider margin than the careful wording suggests. And the most informative documents in the launch are Mistral's own bar charts, which keep showing competitors that the prose around them leaves out.
What exists today
Not much that a machine can read. The launch post calls this a "public
preview" and says the weights "drop end of this month", along with "further details on the model architecture,
additional benchmarks, and our post-training methodology". The Hugging Face API lists no Large 4 repository under
mistralai; the newest thing there is an NVFP4 refresh of Large 3 from September. vLLM's main (commit e82b800,
6 October) has mistral_large_3.py and no Large 4 model file, and mistral-common has no new tokenizer.
So the architecture is three numbers, all from the docs model card: "a granular Mixture-of-Experts architecture. It features 49B active parameters and 1.05T total parameters, and a 1.6B vision encoder." The same card gives a 1M-token context. The blog adds that it was "trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs" in Mistral's own European datacenters, on data spanning more than 160 languages, and that the preview is served on that same hardware.
There is no token count, so the usual 6ND estimate of pretraining compute is off the table. There is no layer count, no expert count, no attention design and no licence name; the docs say "Open" and nothing more.
The predecessor is the best evidence
When a model's own files are missing, the next best thing is its parent's. Mistral Large
3 shipped in December 2025 with weights,
params.json and an Apache 2.0 licence, and the same "granular Mixture-of-Experts" phrase in its card. Its vLLM
implementation is a single short file, and line 12 says everything:
# vllm/model_executor/models/mistral_large_3.py:12
class MistralLarge3ForCausalLM(DeepseekV3ForCausalLM):The other 80 lines are a regex table that renames Mistral's tensors (attention.wq_a, feed_forward.w1,
experts.N.w3) to DeepSeek's (self_attn.q_a_proj, mlp.gate_proj, mlp.experts.N.up_proj). Large 3 is a DeepSeek-V3
block: multi-head latent attention with a 512-wide compressed KV, 61 layers of width 7,168, the first three dense, and
then 128 routed experts of hidden size 4,096 plus one shared expert per layer, four routed experts per token. That is
a reasonable design to copy. I bring it up because it lets me check my own arithmetic before I apply it to a
model nobody can inspect.
Counting layer by layer from params.json, I get 673.42B for the language model and 2.50B for
the Pixtral-style vision tower (48 layers, width 1,664), 675.92B in all. The BF16 repo's index file says
"total_size": 1351982353920, which at two bytes per parameter is 675.99B. The difference, 0.07B, is about the size
of the patch merger and adapter I did not model. The README's own split is "673B params and 39B active" plus "a 2.5B
Vision Encoder", and my active count for the language model is 39.95B with both embedding tables included, 39.01B
without the input one.
That last fact is what "1.05T total, 49B active" means. Large 4 stores 1.55 times what Large 3 did and runs 1.2 times as much per token. The active share drops from 6.1% (41 of 675) to 4.7% (49 of 1,050). Mistral went sparser.
Here is one way to land there, to show the arithmetic rather than to guess the config. Keep Large 3's block exactly and only add experts: each routed expert is 88.1M parameters per layer across 58 MoE layers, so 1.05T needs about 201 experts per layer. With four active that still runs 40B per token; each extra expert per token adds 5.1B, so six per token gives about 50B. Real designs move other dials too (Kimi K3 went to 896 tiny experts with 16 active), and the 1.6B vision encoder alone tells me this is not simply Large 3 widened, because Large 3's tower was 2.5B. A new and smaller encoder fits the "natively multimodal" line: an encoder trained with the model, not bolted on afterwards.
How that sparsity compares with the open models this site has covered, all as each lab reports them:
| model | total | active | active share |
|---|---|---|---|
| Hunyuan Hy3 | 295B | 21B | 7.1% |
| Mistral Large 3 | 675B | 41B | 6.1% |
| GLM-5.3 | 753B | 40B | 5.3% |
| Mistral Large 4 | 1.05T | 49B | 4.7% |
| Reflection Beam | 501B | 23B | 4.6% |
| MiMo-V2.6-Pro | 1,024B | 41.9B | 4.1% |
| Qwen3.8 2.4T | 2.4T | 95B | 4.0% |
| Kimi K3 | 2.8T | 104B | 3.7% |
| LongCat 2.0 | 1.6T | 48B | 3.0% |
Large 4 sits in the middle of the pack, closest in shape to MiMo-V2.6-Pro: a trillion stored, a bit over 40B awake. The mixture-of-experts explainer covers why everyone is converging here.
"Aggregated benchmarks" means Artificial Analysis
The X post does not name the aggregate. The blog quotes Artificial Analysis for most of its coding and cyber numbers, and Artificial Analysis already has a model page for "Mistral Large 4 Preview", so that is the obvious candidate. Its Intelligence Index (v4.3.2) averages ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR.
Large 4 scores 38.4. I pulled every model's record from the page's own data and filtered to the open-weight ones made in the US or Europe. The best is Thinking Machines' Inkling Small at 25.7, then Inkling at 25.0 and NVIDIA's Nemotron 3 Ultra at 22.9. Reflection's Beam, announced yesterday, has no entry yet. So the claim holds by 12.7 points, which is not a narrow win. Mistral's own previous flagship, Large 3, scores 9.3 on the same index, and Medium 3.5 scores 14.2. Whatever else is true, this is a fourfold jump for Mistral in ten months.
Then I sorted the same list without the geography filter, and Large 4 lands eighth.
Above it: MiMo-V2.6-Pro 46.3, GLM-5.3 44.8, Kimi K3 43.6, GLM-5.3 Flash 41.8, Qwen3.8 2.4T 39.9, Qwen3.8 Flash Next 39.8 and DeepSeek V4.1 Flash 39.5. Three of those seven are "flash" models: GLM-5.3 Flash runs 18B of 320B, DeepSeek V4.1 Flash 16B of 552B, and Qwen3.8 Flash Next 6B of 180B. A model with an eighth of Large 4's active parameters scores higher on the index Mistral is citing.
The blog's phrasing is careful about this. It says Large 4 is "competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe". The second half is exactly right. "Competitive" is doing a lot of work in the first: 7.9 points behind the top open model is roughly the gap between Kimi K3 and DeepSeek V4 Pro.
A few more things on the Artificial Analysis page are worth knowing before you use the preview:
- It lists Large 4 as proprietary until the weights are out, which is correct and means it does not appear in their open-weights charts yet. Every comparison above adds it by hand.
- It lists a context window of 524,288 tokens. The docs say 1M. Either the preview endpoint is capped at half, or one of the two pages is wrong; I can't tell which from outside.
- It is verbose. Running the index took 200M output tokens against a median of 81M; GLM-5.3 used 209M and MiMo-V2.6-Pro 144M. At list price that cost $1,602, compared with $2,503 for GLM-5.3, $3,658 for Kimi K3 and $207 for MiMo-V2.6-Pro. Output speed was 116 tokens per second.
- On AA-Omniscience it answers 25.8% of factual questions correctly. Its hallucination rate there, the share of questions it does not know where it answers wrongly instead of declining, is 41.9%, against 29.6% for GLM-5.3 and 53.2% for Kimi K3.
The charts tell a better story than the captions
The launch blog has about twenty charts. I went through each one and wrote down every bar. Most of them include the competitors that beat Large 4, which is to Mistral's credit; the problem is the sentences above them. The widget below puts each claim next to the chart it sits on. The first tab is the aggregate from the previous section.
Claim: “the best open weights model from US or Europe on aggregated benchmarks” (launch post on X)
On the chart: Holds, by almost 13 points over the next Western open model. It is also eighth among the open models Artificial Analysis lists, behind seven Chinese ones, three of them with 18B active parameters or fewer.
Coding is where the prose and the charts agree most. The text says 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, and that a combined Coding Agent Index of 49.8% puts it "ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max". All true, and the comparison is chosen with care: Kimi K3 is ahead on DeepSWE (68 to 62), GLM-5.3 is far ahead on Terminal-Bench (40 to 28), and on SWE-Atlas-QnA Large 4 ties GLM-5.3 at 59 and trails Qwen3.8 Max, Kimi K3 and DeepSeek V4 Pro. The charts' footnote says these are scores "evaluated privately by Artificial Analysis ahead of the harness' public launch". The public board already disagrees slightly: it has Large 4 at 26.8% on Terminal-Bench 4.0, not 28.3%.

That y-axis is the other habit to watch. DeepSWE starts at 40, SciCode-Verified at 65, Finance Agent at 20, Cybench at 60. On the SciCode chart, Large 4's 91.8 against Medium 3.5's 70 is drawn as a bar roughly five times taller, for a model that is 31% better. The widget's "redraw from zero" button puts the same numbers on an honest axis.
The agentic paragraph says Large 4's 59.9% on AutomationBench is "ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro". It is. The chart directly below that sentence also has GLM-5.3 at 62.2.

The science section has the clearest case of a sentence its own chart contradicts. The text: "ML4 is state of the art on SciCode-Verified among open-weight models." The chart: Large 4 91.8, MiMo-V2.6-Pro 91.9, GLM-5.3 92.5, Qwen3.8 2.4T 93.8. All three have public weights. The gaps are small, and SciCode-Verified with six samples per problem is a noisy benchmark at the top, so I would call this a four-way tie. "State of the art" it is not.

Finance is closer than the X post implies. On Vals.ai's Finance Agent v2, Large 4 scores 54.7, which does beat GPT-6 Astra's 53.5, the comparison the text makes. GLM-5.3 is at 55.8 on the same chart. On Finch (FinWorkBench), the spreadsheet benchmark, Kimi K3 is at 77.3 and Large 4 ties DeepSeek V4 Pro at 67.4.

Law is the clean win. On Harvey's Legal Agent benchmark, run by Vals.ai, Large 4 scores 15.8 against Kimi K3's 12.9, Qwen3.8 Max's 10.4 and GPT-6 Astra's 5.4. Every model on that chart is under 16%, so this is a hard benchmark where everyone is bad and Large 4 is less bad, but the lead is real, and unlike most of the post it does not depend on which competitors made the chart.
And "manufacturing", from the X post, has no benchmark behind it at all. The closest thing is an internal human evaluation against GLM-5.3 alone, where expert annotators preferred Large 4 in 62% of CAD comparisons and 68% of STEM ones, 50% in finance and 48% in code. That is a reasonable signal for one pairing. It is not a state-of-the-art claim.
Cyber is where it is different
Everything above is a strong open model doing roughly what the other strong open models do. The cyber section is the part I would actually buy this for, and it is also the part with the most moving pieces.
Three results. On the Artificial Analysis Cyber Index, Large 4 has 50 successes. On CyberGym-E2E, which asks a model to reproduce a real vulnerability in open-source software and then patch it, it scores 82, the top of the chart, with MiMo-V2.6-Pro at 79. On Cybench, 40 capture-the-flag tasks, it solves 93%; on 40 tasks that is 37 solved (92.5%) rounded up, or an average over several runs.

That chart carries the argument Mistral is making. Several models lose much of their cyber score to refusals: Qwen3.8 2.4T has 13 successes and 63 blocks, Opus 5.5 has 29 and 36, GPT-6 Astra 33 and 38. Mistral's point is that defensive security starts with proving a flaw is real, which looks exactly like offence to a refusal filter, and that a defender who loses access mid-incident has a problem. I think that is a fair argument, and it is the strongest case in the launch for open weights specifically, since a self-hosted model is one whose policy you set.
Two caveats. First, the blog says Opus 5.5 and GPT-6 Astra "score near zero" on the CyberGym reproduction test because they refuse. Neither appears on the CyberGym chart, so I can't check that. Second, Mistral published a second version of the Cyber Index chart with a different comparison set, and on that one GLM-5.3 Flash is also at 50. "Among the top five models globally" and "leads open-weight models developed outside China by a wide margin" are both consistent with that. "Leads open-weight models" without the qualifier would not be.

Now the odd part. The safety section reports that Large 4 refuses harmful cyber requests from JailbreakBench, StrongREJECT and AgentHarm 95.3% of the time, higher than Kimi K3 (93.7), GLM-5.3 (87) and DeepSeek V4 Pro (77.7). So it refuses the most on one chart and shows no safety blocks at all on the other. Both can be true, because the prompts are different in kind: the jailbreak sets ask for harm in the open, while the cyber evals frame the work as authorised reproduction inside a sandbox. Whether the line falls in the right place is something a red team finds out, not a bar chart. And the blog says plainly that partners in the red-teaming programme get "the same model with reduced moderation and expanded cyber capabilities". The model you get on the API today and the one cybersecurity partners get are configured differently. Which configuration the cyber charts measured, the post does not say.
Visual grounding, and one half-point
"Surpasses closed frontier models on visual grounding" rests on Dense200, a benchmark of drawing boxes around objects in crowded scenes. Large 4 scores 42.0. GPT-6 Astra scores 41.5. That is the entire comparison with a closed model: one benchmark, one closed competitor, half a point. The open comparisons are more convincing, because Kimi K3 is at 28.9 and DeepSeek V4.1 Flash at 3.3, so whatever Mistral did to the vision side, it works for localisation in a way the other open models' encoders don't.

Elsewhere in multimodal the picture is ordinary. On ChartQA Pro, GPT-6 Astra is ahead (65.1 to 63.1). On GDP.pdf, Kimi K3 is ahead (22 to 18.6). Artificial Analysis has it at 76.4% on MMMU-Pro, below Kimi K3's 80.5%. A 1.6B encoder trained with the model seems to buy fine-grained localisation more than general visual reasoning, which is a plausible thing for a lab selling into manufacturing and earth observation to optimise for.
The RL run is still running
The most technical section of the blog is about post-training, and it has numbers I can do arithmetic on. "At our current scale (3k GPUs), a single training run produces roughly 33 billion tokens per day, of which around 16 billion are trainable completion tokens after filtering and masking." Spread over 3,000 GPUs and 86,400 seconds, 33B tokens a day is about 127 tokens per second per GPU, end to end, including the share of the fleet doing the training rather than the generating. Under half of what gets generated, 48%, survives to become gradient signal; the rest is filtered or masked out. The blog does not break that down further.
The design described is the asynchronous kind every large lab now runs: an autoscaling actor fleet generating rollouts while the trainer updates in parallel, "rollout budgets of millions of tokens across multiple compactions", with care taken to keep the generating policy from drifting too far from the trained one. One composable interface covers chat, science problems, safety, factuality and long tool-use tasks, with reward models, unit tests, LLM judges and static checks combined per task. None of it is unusual. What is unusual is the claim attached: "The reinforcement learning run behind this preview is still in flight, and the model is showing no signs of saturation."

The reward curves only half support that. Customer-support tasks and factuality QA have visibly flattened. Spreadsheet editing and cross-application workflows are noisy and still going up at the right edge. The second chart is better evidence, because it is held-out evaluation rather than training reward:

On the hard coding-agent task, the curve is flat-ish through SFT and climbs steadily through RL to 68.84%. Codebase understanding dips at the start of RL before climbing to 38.66%. Professional tasks reach 1,511 on what looks like an Elo scale. Note that the x-axis is two axes glued together, SFT steps up to 16,388 and then RL steps up to 350, so the visual slope after the dashed line says nothing about efficiency per step. It does say the last RL checkpoint is the best, or level with the best, in all three panels. Codebase understanding has already gone flat over its last two points. That is about as much as "no signs of saturation" can honestly mean.
The practical consequence: the model that ships at the end of October is probably not the one being benchmarked today. Every number in this article is a preview number.
What it takes to run
This is where the absence of a config file matters least, because memory depends mostly on the parameter count, and Mistral published that. 1.05T parameters at one byte each is 1,050 GB in FP8. In BF16 it is 2,100 GB.
For a 4-bit estimate I calibrated on Large 3 rather than using the textbook number. Its NVFP4 repository occupies 403.17 GB for 675.99B parameters, which is 4.77 bits per parameter, not the 4.5 that FP4 plus one FP8 scale per 16 values would give, because the quantization config in that repo leaves the embeddings, the output head, the router gates, the vision tower and every attention projection at higher precision. At 4.77 bits, Large 4 is about 626 GB.
Taking the calculator through the obvious cases:
- FP8 on 8 x H200 (1,128 GB): 78 GB left after weights. That is a working deployment for short contexts and small batches, and very little else.
- FP8 on 8 x B200 (1,440 GB): 390 GB left. This is the realistic single-node target, as it was for Large 3.
- NVFP4 on 8 x H100 (640 GB): the 626 GB of weights fit with 14 GB to spare, which is not enough to serve anything. Large 3's NVFP4 card advertised a single H100 node; Large 4 at the same format does not fit on one.
- A 512 GB Mac Studio needs everything below about 3.9 bits per weight. Two DGX Sparks, which one reply to the launch was hoping to use, need about 1.95.
The cache line in the calculator is a stand-in, labelled as one. If Large 4 keeps Large 3's MLA cache of 576 values per token per layer over 61 layers, a token costs 70,272 bytes in BF16, and a full 1M-token sequence costs about 74 GB. MLA is why a 1M context is affordable at all on this class of model; a plain grouped-query cache would be several times larger. If Mistral changed the attention, this number changes with it.
Per token, the compute is the active count: about 98 GFLOPs for a forward pass at 49B active, roughly a fifth more than Large 3's 41B, and the price follows the active count rather than the stored one. The docs list $1.36 per million input tokens and $4.18 per million output, struck through and halved to $0.68 and $2.09 for the preview, with cached input at $0.07. For comparison, Artificial Analysis lists GLM-5.3 at $1.40 and $4.40, DeepSeek V4 Pro at $1.32 and $3.96, and Kimi K3 at $3 and $15.
What I would look for when the weights land
The first file I will open is params.json. Four numbers in it settle most of what this article had to leave open:
the routed expert count and top-k, which say how Mistral spent the extra 375B parameters; whether kv_lora_rank is
still there, which says whether the attention and its cache carried over from Large 3; and the real context limit,
which settles the 524K against 1M question. The second is the licence. Large 3 was Apache 2.0, and the docs page for
Large 4 says only "Open".
After that, the honest benchmarks will come from other people. Artificial Analysis will move it to the open-weights board and re-run the index on the release checkpoint; the gap to MiMo-V2.6-Pro and GLM-5.3 is the number to watch. The cyber results are the ones I most want reproduced outside Mistral, on the public checkpoint, with the moderation settings stated.
My read today: this is the first open model from Europe that belongs in the same conversation as the Chinese frontier, and it is not yet at the front of that conversation. The cyber behaviour is the distinctive part, and it is the part where "open weights" matters most. The rest of the launch is a strong model described a little more generously than its own charts describe it.
How I checked
- Launch materials: read the X post and its thread through the fxtwitter mirror (the thread's only follow-up is the blog link), the launch blog, and the docs model card, including the page's pricing data, which carries the list price and the preview price as separate fields. Downloaded every chart image the blog serves and read each bar label by eye; every competitor number in this article and in the widget comes from those images.
- Artificial Analysis: parsed the model records embedded in the public "Mistral Large 4 Preview" page on 6
October: intelligence index, per-evaluation scores, licence category, country of the creator, token use, cost and
speed, for every model listed. "Best from the US or Europe" is the highest index among records marked open weights
with creator country
usor a European country. - What is not public: queried the Hugging Face API for
mistralairepositories; checked vLLMmainate82b800andmistral-commonat81708fdfor any Large 4 code. None found. - Large 3 as a reference: read
params.jsonfrommistralai/Mistral-Large-3-675B-Instruct-2512, counted parameters analytically (MLA attention, dense layers, routed and shared experts, both embedding tables, the vision tower) and compared with the BF16 repo'sconsolidated.safetensors.index.jsontotal_size. Readvllm/model_executor/models/mistral_large_3.pyfor the class lineage. The NVFP4 bits-per-parameter figure is that repo's storage over 675.99B parameters. - Arithmetic: memory is parameters times bits over eight; the KV figure is 576 x 61 x 2 bytes per token, from Large 3's config, applied to Large 4 as an assumption. Throughput per GPU is 33B / 3,000 / 86,400.
- Not checked: anything requiring the weights, meaning the architecture, the licence and any benchmark rerun. The claims that closed models score near zero on CyberGym-E2E, the 49.8% Coding Agent Index, the human evaluations, and the "160 languages" figure are Mistral's alone.