# Mistral Large 4: first among Western open models by 13 points, eighth overall

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mistral-large-4
> date: 2026-10-06
> tags: mixture-of-experts, multimodal, benchmarks, security

For about a year, "Le Chonk" was a joke people made at Mistral: a fat cat standing in for the big model everyone
assumed Mistral was not going to ship. This morning [Mistral shipped it](https://x.com/MistralAI/status/2107457414387622310)
under that name. The post makes four claims in four lines: "1T parameters, natively multimodal. 49B active"; "the best
open weights model from US or Europe on aggregated benchmarks"; "state-of-the-art on critical workloads, including
cyber defense, manufacturing and finance"; and that it "surpasses closed frontier models on visual grounding". The
API is live today. "Open weights release end of October."

I wanted to know two things. Is the "best from US or Europe" line true, and which aggregate is it talking about? And
since there are no weights, how much of the rest can anyone check today?

The short version surprised me twice. The headline claim holds, and by a wider margin than the careful wording
suggests. And the most informative documents in the launch are Mistral's own bar charts, which keep showing
competitors that the prose around them leaves out.

## What exists today

Not much that a machine can read. The [launch post](https://mistral.ai/news/mistral-large-4/) calls this a "public
preview" and says the weights "drop end of this month", along with "further details on the model architecture,
additional benchmarks, and our post-training methodology". The Hugging Face API lists no Large 4 repository under
`mistralai`; the newest thing there is an NVFP4 refresh of Large 3 from September. vLLM's `main` (commit `e82b800`,
6 October) has `mistral_large_3.py` and no Large 4 model file, and `mistral-common` has no new tokenizer.

So the architecture is three numbers, all from the [docs model card](https://docs.mistral.ai/models/mistral-large-4):
"a granular Mixture-of-Experts architecture. It features 49B active parameters and 1.05T total parameters, and a 1.6B
vision encoder." The same card gives a 1M-token context. The blog adds that it was "trained from scratch on 3,800
NVIDIA Grace Blackwell GPUs" in Mistral's own European datacenters, on data spanning more than 160 languages, and that
the preview is served on that same hardware.

There is no token count, so the usual 6ND estimate of pretraining compute is off the table. There is no layer
count, no expert count, no attention design and no licence name; the docs say "Open" and nothing more.

## The predecessor is the best evidence

When a model's own files are missing, the next best thing is its parent's. [Mistral Large
3](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512) shipped in December 2025 with weights,
`params.json` and an Apache 2.0 licence, and the same "granular Mixture-of-Experts" phrase in its card. Its vLLM
implementation is a single short file, and line 12 says everything:

```python
# vllm/model_executor/models/mistral_large_3.py:12
class MistralLarge3ForCausalLM(DeepseekV3ForCausalLM):
```

The other 80 lines are a regex table that renames Mistral's tensors (`attention.wq_a`, `feed_forward.w1`,
`experts.N.w3`) to DeepSeek's (`self_attn.q_a_proj`, `mlp.gate_proj`, `mlp.experts.N.up_proj`). Large 3 is a DeepSeek-V3
block: multi-head latent attention with a 512-wide compressed KV, 61 layers of width 7,168, the first three dense, and
then 128 routed experts of hidden size 4,096 plus one shared expert per layer, four routed experts per token. That is
a reasonable design to copy. I bring it up because it lets me check my own arithmetic before I apply it to a
model nobody can inspect.

Counting layer by layer from `params.json`, I get 673.42B for the language model and 2.50B for
the Pixtral-style vision tower (48 layers, width 1,664), 675.92B in all. The BF16 repo's index file says
`"total_size": 1351982353920`, which at two bytes per parameter is 675.99B. The difference, 0.07B, is about the size
of the patch merger and adapter I did not model. The README's own split is "673B params and 39B active" plus "a 2.5B
Vision Encoder", and my active count for the language model is 39.95B with both embedding tables included, 39.01B
without the input one.

<Callout type="note">
Where Large 3's parameters live: 653.9B of its 673.4B language-model parameters, 97.1%, are routed experts. Attention
across all 61 layers is 11.4B. Almost everything a trillion-parameter MoE stores is expert weight that any given token
never touches.
</Callout>

That last fact is what "1.05T total, 49B active" means. Large 4 stores 1.55 times what Large 3 did and runs 1.2 times
as much per token. The active share drops from 6.1% (41 of 675) to 4.7% (49 of 1,050). Mistral went sparser.

Here is one way to land there, to show the arithmetic rather than to guess the config. Keep Large 3's block exactly
and only add experts: each routed expert is 88.1M parameters per layer across 58 MoE layers, so 1.05T needs about 201
experts per layer. With four active that still runs 40B per token; each extra expert per token adds 5.1B, so six per
token gives about 50B. Real designs move other dials too (Kimi K3 went to 896 tiny experts with 16 active), and the
1.6B vision encoder alone tells me this is not simply Large 3 widened, because Large 3's tower was 2.5B. A new and
smaller encoder fits the "natively multimodal" line: an encoder trained with the model, not bolted on afterwards.

How that sparsity compares with the open models this site has covered, all as each lab reports them:

| model | total | active | active share |
|---|---|---|---|
| [Hunyuan Hy3](/articles/hunyuan-hy3) | 295B | 21B | 7.1% |
| Mistral Large 3 | 675B | 41B | 6.1% |
| [GLM-5.3](/articles/glm-5-3) | 753B | 40B | 5.3% |
| **Mistral Large 4** | **1.05T** | **49B** | **4.7%** |
| [Reflection Beam](/articles/reflection-beam) | 501B | 23B | 4.6% |
| [MiMo-V2.6-Pro](/articles/mimo-v2-6) | 1,024B | 41.9B | 4.1% |
| [Qwen3.8 2.4T](/articles/qwen3-8-open-weights) | 2.4T | 95B | 4.0% |
| [Kimi K3](/articles/kimi-k3) | 2.8T | 104B | 3.7% |
| [LongCat 2.0](/articles/longcat-2) | 1.6T | 48B | 3.0% |

Large 4 sits in the middle of the pack, closest in shape to MiMo-V2.6-Pro: a trillion stored, a bit over 40B awake.
The [mixture-of-experts explainer](/architectures/mixture-of-experts) covers why everyone is converging here.

## "Aggregated benchmarks" means Artificial Analysis

The X post does not name the aggregate. The blog quotes Artificial Analysis for most of its coding and cyber numbers,
and Artificial Analysis already has a model page for "Mistral Large 4 Preview", so that is the obvious candidate. Its
Intelligence Index (v4.3.2) averages ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0,
SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR.

Large 4 scores 38.4. I pulled every model's record from the page's own data and filtered to the open-weight ones made
in the US or Europe. The best is Thinking Machines' [Inkling Small](/articles/inkling-small) at 25.7, then
[Inkling](/articles/inkling) at 25.0 and NVIDIA's Nemotron 3 Ultra at 22.9. Reflection's Beam, announced yesterday,
has no entry yet. So the claim holds by 12.7 points, which is not a narrow win. Mistral's own previous flagship, Large
3, scores 9.3 on the same index, and Medium 3.5 scores 14.2. Whatever else is true, this is a fourfold jump for Mistral
in ten months.

Then I sorted the same list without the geography filter, and Large 4 lands eighth.

Above it: MiMo-V2.6-Pro 46.3, GLM-5.3 44.8, Kimi K3 43.6, [GLM-5.3 Flash](/articles/glm-5-3-flash) 41.8, Qwen3.8 2.4T
39.9, [Qwen3.8 Flash Next](/articles/qwen3-8-flash-next) 39.8 and DeepSeek V4.1 Flash 39.5. Three of those seven are
"flash" models: GLM-5.3 Flash runs 18B of 320B, DeepSeek V4.1 Flash 16B of 552B, and Qwen3.8 Flash Next 6B of 180B.
A model with an eighth of Large 4's active parameters scores higher on the index Mistral is citing.

The blog's phrasing is careful about this. It says Large 4 is "competitive with the strongest open-source models
globally, while significantly outperforming any open-weight model developed in the US or Europe". The second half is
exactly right. "Competitive" is doing a lot of work in the first: 7.9 points behind the top open model is roughly the
gap between Kimi K3 and DeepSeek V4 Pro.

A few more things on the Artificial Analysis page are worth knowing before you use the preview:

- It lists Large 4 as proprietary until the weights are out, which is correct and means it does not appear in
  their open-weights charts yet. Every comparison above adds it by hand.
- It lists a context window of 524,288 tokens. The docs say 1M. Either the preview endpoint is capped at half, or one
  of the two pages is wrong; I can't tell which from outside.
- It is verbose. Running the index took 200M output tokens against a median of 81M; GLM-5.3 used 209M and MiMo-V2.6-Pro
  144M. At list price that cost \$1,602, compared with \$2,503 for GLM-5.3, \$3,658 for Kimi K3 and \$207 for
  MiMo-V2.6-Pro. Output speed was 116 tokens per second.
- On AA-Omniscience it answers 25.8% of factual questions correctly. Its hallucination rate there, the share of
  questions it does not know where it answers wrongly instead of declining, is 41.9%, against 29.6% for GLM-5.3 and
  53.2% for Kimi K3.

## The charts tell a better story than the captions

The launch blog has about twenty charts. I went through each one and wrote down every bar. Most of them include the
competitors that beat Large 4, which is to Mistral's credit; the problem is the sentences above them. The widget below
puts each claim next to the chart it sits on. The first tab is the aggregate from the previous section.

<ClaimCheck />

Coding is where the prose and the charts agree most. The text says 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and
28.3% on Terminal-Bench 4, and that a combined Coding Agent Index of 49.8% puts it "ahead of DeepSeek V4 Pro 0813 and
Qwen3.8 Max". All true, and the comparison is chosen with care: Kimi K3 is ahead on DeepSWE (68 to 62), GLM-5.3 is
far ahead on Terminal-Bench (40 to 28), and on SWE-Atlas-QnA Large 4 ties GLM-5.3 at 59 and trails Qwen3.8 Max, Kimi
K3 and DeepSeek V4 Pro. The charts' footnote says these are scores "evaluated privately by Artificial Analysis ahead of
the harness' public launch". The public board already disagrees slightly: it has Large 4 at 26.8% on Terminal-Bench
4.0, not 28.3%.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig1.png"
  alt="Bar chart titled Artificial Analysis DeepSWE 1.1. Mistral Large 4 Preview in orange at 62; Beam (self-reported) 44; Qwen3.8 Max with Claude Code 51; DeepSeek-V4-Pro-0813 with Codex 57; GLM-5.3 with Opencode 61; Kimi K3 with Kimi Code CLI 68. The y-axis runs from 40 to 72."
  caption="DeepSWE 1.1, from a private Artificial Analysis run. Large 4 is second; Kimi K3 leads at 68. The y-axis starts at 40, which makes Beam's 44 look like a sliver (Mistral, Large 4 launch blog, coding chart)."
/>

That y-axis is the other habit to watch. DeepSWE starts at 40, SciCode-Verified at 65, Finance Agent at 20, Cybench at
60. On the SciCode chart, Large 4's 91.8 against Medium 3.5's 70 is drawn as a bar roughly five times taller, for a
model that is 31% better. The widget's "redraw from zero" button puts the same numbers on an honest axis.

The agentic paragraph says Large 4's 59.9% on AutomationBench is "ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4
Pro". It is. The chart directly below that sentence also has GLM-5.3 at 62.2.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig9.png"
  alt="Bar chart titled Artificial Analysis AutomationBench. Mistral Large 4 Preview 59.9; Mistral Medium 3.5 6.3; GLM-5.2 28.4; DeepSeek-V4-Pro-0813 56.7; Qwen3.8 2.4T A95B 57.2; Kimi K3 58.3; GLM-5.3 62.2."
  caption="AutomationBench: Large 4 at 59.9, GLM-5.3 at 62.2 on the same chart. Medium 3.5 at 6.3 is the measure of how far Mistral moved (Mistral, Large 4 launch blog, agentic chart)."
/>

The science section has the clearest case of a sentence its own chart contradicts. The text: "ML4 is state of the art on
SciCode-Verified among open-weight models." The chart: Large 4 91.8, MiMo-V2.6-Pro 91.9, GLM-5.3 92.5, Qwen3.8 2.4T
93.8. All three have public weights. The gaps are small, and SciCode-Verified with six samples per problem is a noisy
benchmark at the top, so I would call this a four-way tie. "State of the art" it is not.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig5.png"
  alt="Bar chart titled SciCode-Verified pass@1 (n=6), y-axis from 65 to 100. Mistral Large 4 Preview 91.8; Mistral Medium 3.5 70; DeepSeek-V4.1-Flash 77.9; GLM-5.2 88.1; Kimi K3 90.3; DeepSeek-V4-Pro-0813 91; Claude Opus 5 91.3; MiMo-V2.6-Pro 91.9; GLM-5.3 92.5; Qwen3.8 2.4T A95B 93.8; GPT-6 Astra 94.2."
  caption="SciCode-Verified pass@1 over six samples. Three open-weight models sit above Large 4's 91.8, on the chart published under the sentence calling it state of the art among open weights (Mistral, Large 4 launch blog, science chart)."
/>

Finance is closer than the X post implies. On Vals.ai's Finance Agent v2, Large 4 scores 54.7, which does beat GPT-6
Astra's 53.5, the comparison the text makes. GLM-5.3 is at 55.8 on the same chart. On Finch (FinWorkBench), the
spreadsheet benchmark, Kimi K3 is at 77.3 and Large 4 ties DeepSeek V4 Pro at 67.4.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig6.png"
  alt="Bar chart titled Vals.ai Finance Agent V2, y-axis from 20 to 60. Mistral Large 4 Preview 54.7; Mistral Medium 3.5 32.1; GLM-5.2 49.7; DeepSeek-V4-Pro-0813 50.4; Qwen3.8 3.8 Max 50.6; Kimi K3 53.1; GPT-6 Astra 53.5; GLM-5.3 55.8."
  caption="Finance Agent v2 from Vals.ai. Large 4 beats GPT-6 Astra by 1.2 points and trails GLM-5.3 by 1.1 (Mistral, Large 4 launch blog, knowledge-work chart)."
/>

Law is the clean win. On Harvey's Legal Agent benchmark, run by Vals.ai, Large 4 scores 15.8 against Kimi K3's 12.9,
Qwen3.8 Max's 10.4 and GPT-6 Astra's 5.4. Every model on that chart is under 16%, so this is a hard benchmark where
everyone is bad and Large 4 is less bad, but the lead is real, and unlike most of the post it does not depend on
which competitors made the chart.

And "manufacturing", from the X post, has no benchmark behind it at all. The closest thing is an internal human
evaluation against GLM-5.3 alone, where expert annotators preferred Large 4 in 62% of CAD comparisons and 68% of STEM
ones, 50% in finance and 48% in code. That is a reasonable signal for one pairing. It is not a state-of-the-art claim.

## Cyber is where it is different

Everything above is a strong open model doing roughly what the other strong open models do. The cyber section is the
part I would actually buy this for, and it is also the part with the most moving pieces.

Three results. On the Artificial Analysis Cyber Index, Large 4 has 50 successes. On CyberGym-E2E, which asks a model to
reproduce a real vulnerability in open-source software and then patch it, it scores 82, the top of the chart, with
MiMo-V2.6-Pro at 79. On Cybench, 40 capture-the-flag tasks, it solves 93%; on 40 tasks that is 37 solved (92.5%)
rounded up, or an average over several runs.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig2.png"
  alt="Bar chart titled Artificial Analysis Cyber Index with a legend for success and safety blocks. Successes: Mistral Large 4 Preview 50, Qwen3.8 2.4T A95B 13, Opus 5.5 29, GPT-6 Astra 33, GLM-5.3 36, Kimi K3 41, DeepSeek-V4.1-Flash 41. Grey bars hanging from the top of the chart show safety blocks: Qwen3.8 63, Opus 5.5 36, GPT-6 Astra 38."
  caption="Artificial Analysis Cyber Index. The bars hanging from the top are safety blocks: tasks a model refused. Opus 5.5, GPT-6 Astra and Qwen3.8 2.4T lose a large share of the index that way; Large 4 shows none (Mistral, Large 4 launch blog, cybersecurity chart)."
/>

That chart carries the argument Mistral is making. Several models lose much of their cyber score to refusals:
Qwen3.8 2.4T has 13 successes and 63 blocks, Opus 5.5 has 29 and 36, GPT-6
Astra 33 and 38. Mistral's point is that defensive security starts with proving a flaw is real, which looks exactly
like offence to a refusal filter, and that a defender who loses access mid-incident has a problem. I think that is a
fair argument, and it is the strongest case in the launch for open weights specifically, since a self-hosted model is
one whose policy you set.

Two caveats. First, the blog says Opus 5.5 and GPT-6 Astra "score near zero" on the CyberGym reproduction test because
they refuse. Neither appears on the CyberGym chart, so I can't check that. Second, Mistral published a second version
of the Cyber Index chart with a different comparison set, and on that one GLM-5.3 Flash is also at 50. "Among the top
five models globally" and "leads open-weight models developed outside China by a wide margin" are both consistent with
that. "Leads open-weight models" without the qualifier would not be.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig3.png"
  alt="Bar chart titled Cybersecurity Benchmarks CyberGym-E2E (AA). Mistral Large 4 Preview 82; DeepSeek-V4.1-Flash 23; GLM-5.3 29; Kimi K3 58; GLM-5.3-Flash 74; Grok 4.7 74; MiMo-V2.6-Pro 79."
  caption="CyberGym-E2E: reproduce a real vulnerability, then patch it. Large 4 leads at 82; the closed models the text says score near zero are not on this chart (Mistral, Large 4 launch blog, cybersecurity chart)."
/>

Now the odd part. The safety section reports that Large 4 refuses harmful cyber requests from JailbreakBench,
StrongREJECT and AgentHarm 95.3% of the time, higher than Kimi K3 (93.7), GLM-5.3 (87) and DeepSeek V4 Pro (77.7). So
it refuses the most on one chart and shows no safety blocks at all on the other. Both can be true, because the prompts are
different in kind: the jailbreak sets ask for harm in the open, while the cyber evals frame the work as authorised
reproduction inside a sandbox. Whether the line falls in the right place is something a red team finds out, not a bar
chart. And the blog says plainly that partners in the red-teaming programme get "the same model with reduced
moderation and expanded cyber capabilities". The model you get on the API today and the one cybersecurity partners
get are configured differently. Which configuration the cyber charts measured, the post does not say.

## Visual grounding, and one half-point

"Surpasses closed frontier models on visual grounding" rests on Dense200, a benchmark of drawing boxes around objects
in crowded scenes. Large 4 scores 42.0. GPT-6 Astra scores 41.5. That is the entire comparison with a closed model:
one benchmark, one closed competitor, half a point. The open comparisons are more convincing, because Kimi K3 is at
28.9 and DeepSeek V4.1 Flash at 3.3, so whatever Mistral did to the vision side, it works for localisation in a way
the other open models' encoders don't.

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig4.png"
  alt="Bar chart titled Multimodal Benchmarks Dense200 (bounding box detection). Mistral Large 4 Preview 42.0; DeepSeek-V4.1-Flash 3.3; Kimi K3 28.9; GPT-6 Astra 41.5."
  caption="Dense200 box detection: 42.0 against GPT-6 Astra's 41.5 is the whole basis for 'surpasses closed frontier models'. The gap to the open models is the more interesting result (Mistral, Large 4 launch blog, multimodal chart)."
/>

Elsewhere in multimodal the picture is ordinary. On ChartQA Pro, GPT-6 Astra is ahead (65.1 to 63.1). On GDP.pdf,
Kimi K3 is ahead (22 to 18.6). Artificial Analysis has it at 76.4% on MMMU-Pro, below Kimi K3's 80.5%. A 1.6B encoder
trained with the model seems to buy fine-grained localisation more than general visual reasoning, which is a
plausible thing for a lab selling into manufacturing and earth observation to optimise for.

## The RL run is still running

The most technical section of the blog is about post-training, and it has numbers I can do arithmetic on. "At our
current scale (3k GPUs), a single training run produces roughly 33 billion tokens per day, of which around 16 billion
are trainable completion tokens after filtering and masking." Spread over 3,000 GPUs and 86,400 seconds, 33B tokens a
day is about 127 tokens per second per GPU, end to end, including the share of the fleet doing the training rather
than the generating. Under half of what gets generated, 48%, survives to become gradient signal; the rest is filtered
or masked out. The blog does not break that down further.

The design described is the asynchronous kind every large lab now runs: an autoscaling actor fleet generating rollouts
while the trainer updates in parallel, "rollout budgets of millions of tokens across multiple compactions", with care
taken to keep the generating policy from drifting too far from the trained one. One composable interface covers chat,
science problems, safety, factuality and long tool-use tasks, with reward models, unit tests, LLM judges and static
checks combined per task. None of it is unusual. What is unusual is the claim attached: "The reinforcement learning
run behind this preview is still in flight, and the model is showing no signs of saturation."

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig7.png"
  alt="Four line charts titled Training reward trends, each running from beginning to end of training: customer-support agentic tasks rising from about 0.45 to about 0.93; agentic cross-application business workflows noisy, from about 0.6 to about 0.8; spreadsheet editing noisy, from about 0.46 to about 0.57; factuality question answering rising fast from about 0.25 then flattening near 0.56."
  caption="Training rewards in four RL environments. Customer support and factuality have flattened; spreadsheet editing and cross-application workflows are noisy and still climbing (Mistral, Large 4 launch blog, RL section)."
/>

The reward curves only half support that. Customer-support tasks and factuality QA have visibly flattened. Spreadsheet
editing and cross-application workflows are noisy and still going up at the right edge. The second chart is better
evidence, because it is held-out evaluation rather than training reward:

<Figure
  src="https://ai.thesatyajit.com/articles/mistral-large-4/fig8.png"
  alt="Three line charts, each with an SFT segment in yellow followed by an RL segment in orange. Hard coding agent task climbs during RL to 68.84%. Codebase understanding tasks is noisy during SFT, dips at the start of RL, then climbs to 38.66%. Agentic professional work tasks climbs steadily to 1,511."
  caption="Downstream evaluations across SFT checkpoints (left of the dashed line, x-axis in SFT steps) and RL checkpoints (right, in RL steps). The x-axis changes units at the line, so the slopes on either side are not comparable (Mistral, Large 4 launch blog, RL section)."
/>

On the hard coding-agent task, the curve is flat-ish through SFT and climbs steadily through RL to 68.84%. Codebase
understanding dips at the start of RL before climbing to 38.66%. Professional tasks reach 1,511 on what looks like an
Elo scale. Note that the x-axis is two axes glued together, SFT steps up to 16,388 and then RL steps up to 350, so the
visual slope after the dashed line says nothing about efficiency per step. It does say the last RL checkpoint is
the best, or level with the best, in all three panels. Codebase understanding has already gone flat over its last two
points. That is about as much as "no signs of saturation" can honestly mean.

The practical consequence: the model that ships at the end of October is probably not the one being benchmarked
today. Every number in this article is a preview number.

## What it takes to run

This is where the absence of a config file matters least, because memory depends mostly on the parameter count, and
Mistral published that. 1.05T parameters at one byte each is 1,050 GB in FP8. In BF16 it is 2,100 GB.

For a 4-bit estimate I calibrated on Large 3 rather than using the textbook number. Its NVFP4 repository occupies
403.17 GB for 675.99B parameters, which is 4.77 bits per parameter, not the 4.5 that FP4 plus one FP8 scale per 16
values would give, because the quantization config in that repo leaves the embeddings, the output head, the router gates, the vision
tower and every attention projection at higher precision. At 4.77 bits, Large 4 is about 626 GB.

<ServingMath />

Taking the calculator through the obvious cases:

- FP8 on 8 x H200 (1,128 GB): 78 GB left after weights. That is a working deployment for short contexts and small
  batches, and very little else.
- FP8 on 8 x B200 (1,440 GB): 390 GB left. This is the realistic single-node target, as it was for Large 3.
- NVFP4 on 8 x H100 (640 GB): the 626 GB of weights fit with 14 GB to spare, which is not enough to serve
  anything. Large 3's NVFP4 card advertised a single H100 node; Large 4 at the same format does not fit on one.
- A 512 GB Mac Studio needs everything below about 3.9 bits per weight. Two DGX Sparks, which one reply to the
  launch was hoping to use, need about 1.95.

The cache line in the calculator is a stand-in, labelled as one. If Large 4 keeps Large 3's MLA cache of 576 values per
token per layer over 61 layers, a token costs 70,272 bytes in BF16, and a full 1M-token sequence costs about 74 GB.
MLA is why a 1M context is affordable at all on this class of model; a plain grouped-query cache would be several
times larger. If Mistral changed the attention, this number changes with it.

Per token, the compute is the active count: about 98 GFLOPs for a forward pass at 49B active, roughly a fifth more
than Large 3's 41B, and the price follows the active count rather than the stored one. The docs list \$1.36 per million input tokens and \$4.18 per million
output, struck through and halved to \$0.68 and \$2.09 for the preview, with cached input at \$0.07. For comparison,
Artificial Analysis lists GLM-5.3 at \$1.40 and \$4.40, DeepSeek V4 Pro at \$1.32 and \$3.96, and Kimi K3 at \$3 and
\$15.

## What I would look for when the weights land

The first file I will open is `params.json`. Four numbers in it settle most of what this article had to leave open:
the routed expert count and top-k, which say how Mistral spent the extra 375B parameters; whether `kv_lora_rank` is
still there, which says whether the attention and its cache carried over from Large 3; and the real context limit,
which settles the 524K against 1M question. The second is the licence. Large 3 was Apache 2.0, and the docs page for
Large 4 says only "Open".

After that, the honest benchmarks will come from other people. Artificial Analysis will move it to the open-weights
board and re-run the index on the release checkpoint; the gap to MiMo-V2.6-Pro and GLM-5.3 is the number to watch. The
cyber results are the ones I most want reproduced outside Mistral, on the public checkpoint, with the moderation
settings stated.

My read today: this is the first open model from Europe that belongs in the same conversation as the Chinese
frontier, and it is not yet at the front of that conversation. The cyber behaviour is the distinctive part, and it is
the part where "open weights" matters most. The rest of the launch is a strong model described a little more
generously than its own charts describe it.

## How I checked

- Launch materials: read the X post and its thread through the fxtwitter mirror (the thread's only follow-up is
  the blog link), the launch blog, and the docs model card, including the page's pricing data, which carries the list
  price and the preview price as separate fields. Downloaded every chart image the blog serves and read each bar
  label by eye; every competitor number in this article and in the widget comes from those images.
- Artificial Analysis: parsed the model records embedded in the public "Mistral Large 4 Preview" page on 6
  October: intelligence index, per-evaluation scores, licence category, country of the creator, token use, cost and
  speed, for every model listed. "Best from the US or Europe" is the highest index among records marked open weights
  with creator country `us` or a European country.
- What is not public: queried the Hugging Face API for `mistralai` repositories; checked vLLM `main` at `e82b800`
  and `mistral-common` at `81708fd` for any Large 4 code. None found.
- Large 3 as a reference: read `params.json` from `mistralai/Mistral-Large-3-675B-Instruct-2512`, counted
  parameters analytically (MLA attention, dense layers, routed and shared experts, both embedding tables, the vision
  tower) and compared with the BF16 repo's `consolidated.safetensors.index.json` `total_size`. Read
  `vllm/model_executor/models/mistral_large_3.py` for the class lineage. The NVFP4 bits-per-parameter figure is that
  repo's storage over 675.99B parameters.
- Arithmetic: memory is parameters times bits over eight; the KV figure is 576 x 61 x 2 bytes per token, from
  Large 3's config, applied to Large 4 as an assumption. Throughput per GPU is 33B / 3,000 / 86,400.
- Not checked: anything requiring the weights, meaning the architecture, the licence and any benchmark rerun. The
  claims that closed models score near zero on CyberGym-E2E, the 49.8% Coding Agent Index, the human evaluations, and
  the "160 languages" figure are Mistral's alone.
