~/satyajit

Mistral Large 4: first among Western open models by 13 points, eighth overall

mdjsonmcp

2026-10-06 · 24 min · mixture-of-experts · multimodal · benchmarks · security

Why read this

Notabletop 60%

Checks Large 4's 'best from US or Europe' against Artificial Analysis records (true; eighth overall) and its charts against its prose; plus a memory planner.

  • Checked against the source
  • Widely used
  • Concrete numbers to act on

LLM architectureAPI onlyPractitioner model

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
3 of 3: Reproduces the headline result, or shows from primary files it is wrong
Can I run it?
1 of 3: API-only, gated or restrictive licence
Will I understand it?
1 of 3: Partial mechanism
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
2 of 3: A widely used model, tool or lab release
Only here?
2 of 3: A teardown or measurement few others did

Score 59 of 100, ranked 249 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

For about a year, "Le Chonk" was a joke people made at Mistral: a fat cat standing in for the big model everyone assumed Mistral was not going to ship. This morning Mistral shipped it under that name. The post makes four claims in four lines: "1T parameters, natively multimodal. 49B active"; "the best open weights model from US or Europe on aggregated benchmarks"; "state-of-the-art on critical workloads, including cyber defense, manufacturing and finance"; and that it "surpasses closed frontier models on visual grounding". The API is live today. "Open weights release end of October."

I wanted to know two things. Is the "best from US or Europe" line true, and which aggregate is it talking about? And since there are no weights, how much of the rest can anyone check today?

The short version surprised me twice. The headline claim holds, and by a wider margin than the careful wording suggests. And the most informative documents in the launch are Mistral's own bar charts, which keep showing competitors that the prose around them leaves out.

What exists today

Not much that a machine can read. The launch post calls this a "public preview" and says the weights "drop end of this month", along with "further details on the model architecture, additional benchmarks, and our post-training methodology". The Hugging Face API lists no Large 4 repository under mistralai; the newest thing there is an NVFP4 refresh of Large 3 from September. vLLM's main (commit e82b800, 6 October) has mistral_large_3.py and no Large 4 model file, and mistral-common has no new tokenizer.

So the architecture is three numbers, all from the docs model card: "a granular Mixture-of-Experts architecture. It features 49B active parameters and 1.05T total parameters, and a 1.6B vision encoder." The same card gives a 1M-token context. The blog adds that it was "trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs" in Mistral's own European datacenters, on data spanning more than 160 languages, and that the preview is served on that same hardware.

There is no token count, so the usual 6ND estimate of pretraining compute is off the table. There is no layer count, no expert count, no attention design and no licence name; the docs say "Open" and nothing more.

The predecessor is the best evidence

When a model's own files are missing, the next best thing is its parent's. Mistral Large 3 shipped in December 2025 with weights, params.json and an Apache 2.0 licence, and the same "granular Mixture-of-Experts" phrase in its card. Its vLLM implementation is a single short file, and line 12 says everything:

# vllm/model_executor/models/mistral_large_3.py:12
class MistralLarge3ForCausalLM(DeepseekV3ForCausalLM):

The other 80 lines are a regex table that renames Mistral's tensors (attention.wq_a, feed_forward.w1, experts.N.w3) to DeepSeek's (self_attn.q_a_proj, mlp.gate_proj, mlp.experts.N.up_proj). Large 3 is a DeepSeek-V3 block: multi-head latent attention with a 512-wide compressed KV, 61 layers of width 7,168, the first three dense, and then 128 routed experts of hidden size 4,096 plus one shared expert per layer, four routed experts per token. That is a reasonable design to copy. I bring it up because it lets me check my own arithmetic before I apply it to a model nobody can inspect.

Counting layer by layer from params.json, I get 673.42B for the language model and 2.50B for the Pixtral-style vision tower (48 layers, width 1,664), 675.92B in all. The BF16 repo's index file says "total_size": 1351982353920, which at two bytes per parameter is 675.99B. The difference, 0.07B, is about the size of the patch merger and adapter I did not model. The README's own split is "673B params and 39B active" plus "a 2.5B Vision Encoder", and my active count for the language model is 39.95B with both embedding tables included, 39.01B without the input one.

That last fact is what "1.05T total, 49B active" means. Large 4 stores 1.55 times what Large 3 did and runs 1.2 times as much per token. The active share drops from 6.1% (41 of 675) to 4.7% (49 of 1,050). Mistral went sparser.

Here is one way to land there, to show the arithmetic rather than to guess the config. Keep Large 3's block exactly and only add experts: each routed expert is 88.1M parameters per layer across 58 MoE layers, so 1.05T needs about 201 experts per layer. With four active that still runs 40B per token; each extra expert per token adds 5.1B, so six per token gives about 50B. Real designs move other dials too (Kimi K3 went to 896 tiny experts with 16 active), and the 1.6B vision encoder alone tells me this is not simply Large 3 widened, because Large 3's tower was 2.5B. A new and smaller encoder fits the "natively multimodal" line: an encoder trained with the model, not bolted on afterwards.

How that sparsity compares with the open models this site has covered, all as each lab reports them:

modeltotalactiveactive share
Hunyuan Hy3295B21B7.1%
Mistral Large 3675B41B6.1%
GLM-5.3753B40B5.3%
Mistral Large 41.05T49B4.7%
Reflection Beam501B23B4.6%
MiMo-V2.6-Pro1,024B41.9B4.1%
Qwen3.8 2.4T2.4T95B4.0%
Kimi K32.8T104B3.7%
LongCat 2.01.6T48B3.0%

Large 4 sits in the middle of the pack, closest in shape to MiMo-V2.6-Pro: a trillion stored, a bit over 40B awake. The mixture-of-experts explainer covers why everyone is converging here.

"Aggregated benchmarks" means Artificial Analysis

The X post does not name the aggregate. The blog quotes Artificial Analysis for most of its coding and cyber numbers, and Artificial Analysis already has a model page for "Mistral Large 4 Preview", so that is the obvious candidate. Its Intelligence Index (v4.3.2) averages ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR.

Large 4 scores 38.4. I pulled every model's record from the page's own data and filtered to the open-weight ones made in the US or Europe. The best is Thinking Machines' Inkling Small at 25.7, then Inkling at 25.0 and NVIDIA's Nemotron 3 Ultra at 22.9. Reflection's Beam, announced yesterday, has no entry yet. So the claim holds by 12.7 points, which is not a narrow win. Mistral's own previous flagship, Large 3, scores 9.3 on the same index, and Medium 3.5 scores 14.2. Whatever else is true, this is a fourfold jump for Mistral in ten months.

Then I sorted the same list without the geography filter, and Large 4 lands eighth.

Above it: MiMo-V2.6-Pro 46.3, GLM-5.3 44.8, Kimi K3 43.6, GLM-5.3 Flash 41.8, Qwen3.8 2.4T 39.9, Qwen3.8 Flash Next 39.8 and DeepSeek V4.1 Flash 39.5. Three of those seven are "flash" models: GLM-5.3 Flash runs 18B of 320B, DeepSeek V4.1 Flash 16B of 552B, and Qwen3.8 Flash Next 6B of 180B. A model with an eighth of Large 4's active parameters scores higher on the index Mistral is citing.

The blog's phrasing is careful about this. It says Large 4 is "competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe". The second half is exactly right. "Competitive" is doing a lot of work in the first: 7.9 points behind the top open model is roughly the gap between Kimi K3 and DeepSeek V4 Pro.

A few more things on the Artificial Analysis page are worth knowing before you use the preview:

The charts tell a better story than the captions

The launch blog has about twenty charts. I went through each one and wrote down every bar. Most of them include the competitors that beat Large 4, which is to Mistral's credit; the problem is the sentences above them. The widget below puts each claim next to the chart it sits on. The first tab is the aggregate from the previous section.

The claim, then the chart it sits on

Claim: “the best open weights model from US or Europe on aggregated benchmarks” (launch post on X)

MiMo-V2.6-Pro
46.3
GLM-5.3
44.8
Kimi K3
43.6
GLM-5.3 Flash
41.8320B / 18B
Qwen3.8 2.4T
39.9
Qwen3.8 Flash Next
39.8180B / 6B
DeepSeek V4.1 Flash
39.5552B / 16B
Mistral Large 4
38.4
DeepSeek V4 Pro
36.0
Inkling Small
25.7
Inkling
25.0
Nemotron 3 Ultra
22.9
Mistral Large 3
9.3
axis starts at 0 · rank 8 of 13
Mistral Large 4US or Europe, openChina, openclosed

On the chart: Holds, by almost 13 points over the next Western open model. It is also eighth among the open models Artificial Analysis lists, behind seven Chinese ones, three of them with 18B active parameters or fewer.

Values copied from Mistral's launch charts; the first tab is Artificial Analysis's public Intelligence Index for open-weight models, with Mistral Large 4 added although Artificial Analysis lists it as proprietary until the weights ship. Nothing re-run.

Coding is where the prose and the charts agree most. The text says 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, and that a combined Coding Agent Index of 49.8% puts it "ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max". All true, and the comparison is chosen with care: Kimi K3 is ahead on DeepSWE (68 to 62), GLM-5.3 is far ahead on Terminal-Bench (40 to 28), and on SWE-Atlas-QnA Large 4 ties GLM-5.3 at 59 and trails Qwen3.8 Max, Kimi K3 and DeepSeek V4 Pro. The charts' footnote says these are scores "evaluated privately by Artificial Analysis ahead of the harness' public launch". The public board already disagrees slightly: it has Large 4 at 26.8% on Terminal-Bench 4.0, not 28.3%.

Bar chart titled Artificial Analysis DeepSWE 1.1. Mistral Large 4 Preview in orange at 62; Beam (self-reported) 44; Qwen3.8 Max with Claude Code 51; DeepSeek-V4-Pro-0813 with Codex 57; GLM-5.3 with Opencode 61; Kimi K3 with Kimi Code CLI 68. The y-axis runs from 40 to 72.
DeepSWE 1.1, from a private Artificial Analysis run. Large 4 is second; Kimi K3 leads at 68. The y-axis starts at 40, which makes Beam's 44 look like a sliver (Mistral, Large 4 launch blog, coding chart).

That y-axis is the other habit to watch. DeepSWE starts at 40, SciCode-Verified at 65, Finance Agent at 20, Cybench at 60. On the SciCode chart, Large 4's 91.8 against Medium 3.5's 70 is drawn as a bar roughly five times taller, for a model that is 31% better. The widget's "redraw from zero" button puts the same numbers on an honest axis.

The agentic paragraph says Large 4's 59.9% on AutomationBench is "ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro". It is. The chart directly below that sentence also has GLM-5.3 at 62.2.

Bar chart titled Artificial Analysis AutomationBench. Mistral Large 4 Preview 59.9; Mistral Medium 3.5 6.3; GLM-5.2 28.4; DeepSeek-V4-Pro-0813 56.7; Qwen3.8 2.4T A95B 57.2; Kimi K3 58.3; GLM-5.3 62.2.
AutomationBench: Large 4 at 59.9, GLM-5.3 at 62.2 on the same chart. Medium 3.5 at 6.3 is the measure of how far Mistral moved (Mistral, Large 4 launch blog, agentic chart).

The science section has the clearest case of a sentence its own chart contradicts. The text: "ML4 is state of the art on SciCode-Verified among open-weight models." The chart: Large 4 91.8, MiMo-V2.6-Pro 91.9, GLM-5.3 92.5, Qwen3.8 2.4T 93.8. All three have public weights. The gaps are small, and SciCode-Verified with six samples per problem is a noisy benchmark at the top, so I would call this a four-way tie. "State of the art" it is not.

Bar chart titled SciCode-Verified pass@1 (n=6), y-axis from 65 to 100. Mistral Large 4 Preview 91.8; Mistral Medium 3.5 70; DeepSeek-V4.1-Flash 77.9; GLM-5.2 88.1; Kimi K3 90.3; DeepSeek-V4-Pro-0813 91; Claude Opus 5 91.3; MiMo-V2.6-Pro 91.9; GLM-5.3 92.5; Qwen3.8 2.4T A95B 93.8; GPT-6 Astra 94.2.
SciCode-Verified pass@1 over six samples. Three open-weight models sit above Large 4's 91.8, on the chart published under the sentence calling it state of the art among open weights (Mistral, Large 4 launch blog, science chart).

Finance is closer than the X post implies. On Vals.ai's Finance Agent v2, Large 4 scores 54.7, which does beat GPT-6 Astra's 53.5, the comparison the text makes. GLM-5.3 is at 55.8 on the same chart. On Finch (FinWorkBench), the spreadsheet benchmark, Kimi K3 is at 77.3 and Large 4 ties DeepSeek V4 Pro at 67.4.

Bar chart titled Vals.ai Finance Agent V2, y-axis from 20 to 60. Mistral Large 4 Preview 54.7; Mistral Medium 3.5 32.1; GLM-5.2 49.7; DeepSeek-V4-Pro-0813 50.4; Qwen3.8 3.8 Max 50.6; Kimi K3 53.1; GPT-6 Astra 53.5; GLM-5.3 55.8.
Finance Agent v2 from Vals.ai. Large 4 beats GPT-6 Astra by 1.2 points and trails GLM-5.3 by 1.1 (Mistral, Large 4 launch blog, knowledge-work chart).

Law is the clean win. On Harvey's Legal Agent benchmark, run by Vals.ai, Large 4 scores 15.8 against Kimi K3's 12.9, Qwen3.8 Max's 10.4 and GPT-6 Astra's 5.4. Every model on that chart is under 16%, so this is a hard benchmark where everyone is bad and Large 4 is less bad, but the lead is real, and unlike most of the post it does not depend on which competitors made the chart.

And "manufacturing", from the X post, has no benchmark behind it at all. The closest thing is an internal human evaluation against GLM-5.3 alone, where expert annotators preferred Large 4 in 62% of CAD comparisons and 68% of STEM ones, 50% in finance and 48% in code. That is a reasonable signal for one pairing. It is not a state-of-the-art claim.

Cyber is where it is different

Everything above is a strong open model doing roughly what the other strong open models do. The cyber section is the part I would actually buy this for, and it is also the part with the most moving pieces.

Three results. On the Artificial Analysis Cyber Index, Large 4 has 50 successes. On CyberGym-E2E, which asks a model to reproduce a real vulnerability in open-source software and then patch it, it scores 82, the top of the chart, with MiMo-V2.6-Pro at 79. On Cybench, 40 capture-the-flag tasks, it solves 93%; on 40 tasks that is 37 solved (92.5%) rounded up, or an average over several runs.

Bar chart titled Artificial Analysis Cyber Index with a legend for success and safety blocks. Successes: Mistral Large 4 Preview 50, Qwen3.8 2.4T A95B 13, Opus 5.5 29, GPT-6 Astra 33, GLM-5.3 36, Kimi K3 41, DeepSeek-V4.1-Flash 41. Grey bars hanging from the top of the chart show safety blocks: Qwen3.8 63, Opus 5.5 36, GPT-6 Astra 38.
Artificial Analysis Cyber Index. The bars hanging from the top are safety blocks: tasks a model refused. Opus 5.5, GPT-6 Astra and Qwen3.8 2.4T lose a large share of the index that way; Large 4 shows none (Mistral, Large 4 launch blog, cybersecurity chart).

That chart carries the argument Mistral is making. Several models lose much of their cyber score to refusals: Qwen3.8 2.4T has 13 successes and 63 blocks, Opus 5.5 has 29 and 36, GPT-6 Astra 33 and 38. Mistral's point is that defensive security starts with proving a flaw is real, which looks exactly like offence to a refusal filter, and that a defender who loses access mid-incident has a problem. I think that is a fair argument, and it is the strongest case in the launch for open weights specifically, since a self-hosted model is one whose policy you set.

Two caveats. First, the blog says Opus 5.5 and GPT-6 Astra "score near zero" on the CyberGym reproduction test because they refuse. Neither appears on the CyberGym chart, so I can't check that. Second, Mistral published a second version of the Cyber Index chart with a different comparison set, and on that one GLM-5.3 Flash is also at 50. "Among the top five models globally" and "leads open-weight models developed outside China by a wide margin" are both consistent with that. "Leads open-weight models" without the qualifier would not be.

Bar chart titled Cybersecurity Benchmarks CyberGym-E2E (AA). Mistral Large 4 Preview 82; DeepSeek-V4.1-Flash 23; GLM-5.3 29; Kimi K3 58; GLM-5.3-Flash 74; Grok 4.7 74; MiMo-V2.6-Pro 79.
CyberGym-E2E: reproduce a real vulnerability, then patch it. Large 4 leads at 82; the closed models the text says score near zero are not on this chart (Mistral, Large 4 launch blog, cybersecurity chart).

Now the odd part. The safety section reports that Large 4 refuses harmful cyber requests from JailbreakBench, StrongREJECT and AgentHarm 95.3% of the time, higher than Kimi K3 (93.7), GLM-5.3 (87) and DeepSeek V4 Pro (77.7). So it refuses the most on one chart and shows no safety blocks at all on the other. Both can be true, because the prompts are different in kind: the jailbreak sets ask for harm in the open, while the cyber evals frame the work as authorised reproduction inside a sandbox. Whether the line falls in the right place is something a red team finds out, not a bar chart. And the blog says plainly that partners in the red-teaming programme get "the same model with reduced moderation and expanded cyber capabilities". The model you get on the API today and the one cybersecurity partners get are configured differently. Which configuration the cyber charts measured, the post does not say.

Visual grounding, and one half-point

"Surpasses closed frontier models on visual grounding" rests on Dense200, a benchmark of drawing boxes around objects in crowded scenes. Large 4 scores 42.0. GPT-6 Astra scores 41.5. That is the entire comparison with a closed model: one benchmark, one closed competitor, half a point. The open comparisons are more convincing, because Kimi K3 is at 28.9 and DeepSeek V4.1 Flash at 3.3, so whatever Mistral did to the vision side, it works for localisation in a way the other open models' encoders don't.

Bar chart titled Multimodal Benchmarks Dense200 (bounding box detection). Mistral Large 4 Preview 42.0; DeepSeek-V4.1-Flash 3.3; Kimi K3 28.9; GPT-6 Astra 41.5.
Dense200 box detection: 42.0 against GPT-6 Astra's 41.5 is the whole basis for 'surpasses closed frontier models'. The gap to the open models is the more interesting result (Mistral, Large 4 launch blog, multimodal chart).

Elsewhere in multimodal the picture is ordinary. On ChartQA Pro, GPT-6 Astra is ahead (65.1 to 63.1). On GDP.pdf, Kimi K3 is ahead (22 to 18.6). Artificial Analysis has it at 76.4% on MMMU-Pro, below Kimi K3's 80.5%. A 1.6B encoder trained with the model seems to buy fine-grained localisation more than general visual reasoning, which is a plausible thing for a lab selling into manufacturing and earth observation to optimise for.

The RL run is still running

The most technical section of the blog is about post-training, and it has numbers I can do arithmetic on. "At our current scale (3k GPUs), a single training run produces roughly 33 billion tokens per day, of which around 16 billion are trainable completion tokens after filtering and masking." Spread over 3,000 GPUs and 86,400 seconds, 33B tokens a day is about 127 tokens per second per GPU, end to end, including the share of the fleet doing the training rather than the generating. Under half of what gets generated, 48%, survives to become gradient signal; the rest is filtered or masked out. The blog does not break that down further.

The design described is the asynchronous kind every large lab now runs: an autoscaling actor fleet generating rollouts while the trainer updates in parallel, "rollout budgets of millions of tokens across multiple compactions", with care taken to keep the generating policy from drifting too far from the trained one. One composable interface covers chat, science problems, safety, factuality and long tool-use tasks, with reward models, unit tests, LLM judges and static checks combined per task. None of it is unusual. What is unusual is the claim attached: "The reinforcement learning run behind this preview is still in flight, and the model is showing no signs of saturation."

Four line charts titled Training reward trends, each running from beginning to end of training: customer-support agentic tasks rising from about 0.45 to about 0.93; agentic cross-application business workflows noisy, from about 0.6 to about 0.8; spreadsheet editing noisy, from about 0.46 to about 0.57; factuality question answering rising fast from about 0.25 then flattening near 0.56.
Training rewards in four RL environments. Customer support and factuality have flattened; spreadsheet editing and cross-application workflows are noisy and still climbing (Mistral, Large 4 launch blog, RL section).

The reward curves only half support that. Customer-support tasks and factuality QA have visibly flattened. Spreadsheet editing and cross-application workflows are noisy and still going up at the right edge. The second chart is better evidence, because it is held-out evaluation rather than training reward:

Three line charts, each with an SFT segment in yellow followed by an RL segment in orange. Hard coding agent task climbs during RL to 68.84%. Codebase understanding tasks is noisy during SFT, dips at the start of RL, then climbs to 38.66%. Agentic professional work tasks climbs steadily to 1,511.
Downstream evaluations across SFT checkpoints (left of the dashed line, x-axis in SFT steps) and RL checkpoints (right, in RL steps). The x-axis changes units at the line, so the slopes on either side are not comparable (Mistral, Large 4 launch blog, RL section).

On the hard coding-agent task, the curve is flat-ish through SFT and climbs steadily through RL to 68.84%. Codebase understanding dips at the start of RL before climbing to 38.66%. Professional tasks reach 1,511 on what looks like an Elo scale. Note that the x-axis is two axes glued together, SFT steps up to 16,388 and then RL steps up to 350, so the visual slope after the dashed line says nothing about efficiency per step. It does say the last RL checkpoint is the best, or level with the best, in all three panels. Codebase understanding has already gone flat over its last two points. That is about as much as "no signs of saturation" can honestly mean.

The practical consequence: the model that ships at the end of October is probably not the one being benchmarked today. Every number in this article is a preview number.

What it takes to run

This is where the absence of a config file matters least, because memory depends mostly on the parameter count, and Mistral published that. 1.05T parameters at one byte each is 1,050 GB in FP8. In BF16 it is 2,100 GB.

For a 4-bit estimate I calibrated on Large 3 rather than using the textbook number. Its NVFP4 repository occupies 403.17 GB for 675.99B parameters, which is 4.77 bits per parameter, not the 4.5 that FP4 plus one FP8 scale per 16 values would give, because the quantization config in that repo leaves the embeddings, the output head, the router gates, the vision tower and every attention projection at higher precision. At 4.77 bits, Large 4 is about 626 GB.

Does it fit?
weights format
hardware
weightsKV cache, Large 3 shape| black line = 8 x H200 memory
weights: 1050 GB (8 bits x 1.05T, docs model card)KV cache: 9.2 GB (70,272 B per token)8 x H200: 1128 GBleft for activations and batching: 69 GB
Weight memory is exact arithmetic on the published parameter count. The cache line is a stand-in: Mistral has not published Large 4's attention, so it assumes Large 3's MLA cache. Runtime overhead, activations and the vision encoder's working memory are not counted.

Taking the calculator through the obvious cases:

The cache line in the calculator is a stand-in, labelled as one. If Large 4 keeps Large 3's MLA cache of 576 values per token per layer over 61 layers, a token costs 70,272 bytes in BF16, and a full 1M-token sequence costs about 74 GB. MLA is why a 1M context is affordable at all on this class of model; a plain grouped-query cache would be several times larger. If Mistral changed the attention, this number changes with it.

Per token, the compute is the active count: about 98 GFLOPs for a forward pass at 49B active, roughly a fifth more than Large 3's 41B, and the price follows the active count rather than the stored one. The docs list $1.36 per million input tokens and $4.18 per million output, struck through and halved to $0.68 and $2.09 for the preview, with cached input at $0.07. For comparison, Artificial Analysis lists GLM-5.3 at $1.40 and $4.40, DeepSeek V4 Pro at $1.32 and $3.96, and Kimi K3 at $3 and $15.

What I would look for when the weights land

The first file I will open is params.json. Four numbers in it settle most of what this article had to leave open: the routed expert count and top-k, which say how Mistral spent the extra 375B parameters; whether kv_lora_rank is still there, which says whether the attention and its cache carried over from Large 3; and the real context limit, which settles the 524K against 1M question. The second is the licence. Large 3 was Apache 2.0, and the docs page for Large 4 says only "Open".

After that, the honest benchmarks will come from other people. Artificial Analysis will move it to the open-weights board and re-run the index on the release checkpoint; the gap to MiMo-V2.6-Pro and GLM-5.3 is the number to watch. The cyber results are the ones I most want reproduced outside Mistral, on the public checkpoint, with the moderation settings stated.

My read today: this is the first open model from Europe that belongs in the same conversation as the Chinese frontier, and it is not yet at the front of that conversation. The cyber behaviour is the distinctive part, and it is the part where "open weights" matters most. The rest of the launch is a strong model described a little more generously than its own charts describe it.

How I checked

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Mistral Large 4: first among Western open models by 13 points, eighth overall", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026mistrallarge4,
  author = {Satyajit Ghana},
  title  = {Mistral Large 4: first among Western open models by 13 points, eighth overall},
  url    = {https://ai.thesatyajit.com/articles/mistral-large-4},
  year   = {2026}
}
share