# Mellum2.1: JetBrains' 2.5B-active coding agent, and which speed it is fast at

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mellum2-1
> date: 2026-10-08
> tags: agentic-coding, mixture-of-experts, reinforcement-learning, small-models, kv-cache, benchmarks, gguf

Nikita Pavlichenko's announcement of Mellum2.1 has one sentence that does a lot of work: it "achieves top scores on the agentic coding benchmarks relative to its speed." Relative to its speed is a hedge, and an honest one. It tells you the model does not top the agentic benchmarks outright. So I wanted to know three things. Which speed, measured how, and against whom? Where does the speed come from, given that this is a 12B model? And what did JetBrains actually change between June's Mellum2 and this release?

The short version: the speed is real and it is architectural. Only 8 of 64 experts run per token, and three layers in four only look back 1,024 tokens. The release itself changes none of that; every byte of difference is post-training. The benchmark jump on agentic work is enormous, from 2.0 to 47.0 on SWE-bench Verified. But the speed chart was measured on a 2,304-token-in, 256-token-out workload, which is the shape of a code completion and not of an agent turn. The one rival that beats Mellum2.1 on every agentic benchmark, Qwen3.5-9B, is also the one it is nearly twice as fast as.

<Figure
  src="https://ai.thesatyajit.com/articles/mellum2-1/fig1.png"
  alt="Grouped bar chart titled Mellum2.1, Thinking model for developer workflows. Six benchmarks, each with four bars for Mellum2.1, Mellum2, Qwen3.5-9B and Gemma 4 E4B. Mellum2.1 scores 82.0 on LiveCodeBench, 83.3 on AIME, 64.6 on GPQA Diamond, 62.3 on BFCL v4, 90.6 on IFEval and 47.0 on SWE-bench Verified. Gains over Mellum2 are printed beneath: plus 12.6, 23.2, 13.6, 12.7, 11.1 and 45.0. Mellum2's SWE-bench bar is barely visible."
  caption="JetBrains' summary chart. Every score is self-reported and was run by JetBrains through one pipeline for all four models; the SWE-bench Verified jump of 45.0 points is the headline (JetBrains blog, Mellum2.1 summary chart)."
/>

## What is in the box

There are two repositories: [`JetBrains/Mellum2.1-12B-A2.5B-Thinking`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) with bf16 safetensors, and [`JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF) with five llama.cpp files. Both are Apache 2.0, in the card metadata and inside the GGUF header (`general.license = apache-2.0`).

<ModelCard repo="JetBrains/Mellum2.1-12B-A2.5B-Thinking" note="12,149,923,072 parameters in 5,631 bf16 tensors across 5 shards (measured from the headers). 64 experts per layer, 8 routed per token, no shared expert; 2.44B active per token including both embedding matrices. No MTP tensors in this release. Apache-2.0." />

<ModelCard repo="JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF" note="BF16, Q8_0, Q6_K, Q4_K_M (recommended, 8.07 GB) and MXFP4_MOE (7.03 GB). GGUF architecture string: mellum. Apache-2.0." />

I read the safetensors headers with HTTP range requests rather than downloading 24 GB. There are 5,631 tensors, all bf16, summing to 12,149,923,072 parameters, which matches the `total_parameters` field in `model.safetensors.index.json` exactly. The experts are almost all of it: 84 groups of 64 experts (28 layers, each with gate, up and down projections) come to 11,098,128,384 parameters, or 91.3% of the model. Everything else, attention plus embeddings plus norms plus routers, is about 1.05B.

That split is what the "A2.5B" in the name is counting. Each token goes through every non-expert weight and through 8 of the 64 experts in each layer, so the active count is 1.05B plus one eighth of 11.1B. I get 2,439,060,736, or 2.44B. If you leave out the input embedding table (a lookup, not a matrix multiply) it is 2.21B. JetBrains rounds to 2.5B, which is fair either way.

Then I diffed `config.json` against Mellum2 Thinking's from June. The two are identical, key for key. The blog says "the architecture hasn't changed since version 2," and the config agrees: 28 layers, hidden size 2,304, 32 query heads and 4 KV heads of width 128, a 98,304-token vocabulary, and the same `layer_types` list. That list is the interesting one. It reads `sliding, sliding, sliding, full` seven times over. One oddity: the config still declares an `intermediate_size` of 7,168 for a dense MLP, but every entry in `mlp_layer_types` is `sparse`, and no dense MLP weight exists in the headers. It is a leftover field.

What is missing is the multi-token prediction head. The Mellum2 technical report describes "a single Multi-Token Prediction head that doubles as both an auxiliary pre-training objective and a built-in draft model for speculative decoding," and both the blog and the X thread promise it for this release. No tensor in the headers has `mtp` or `nextn` in its name. It is coming once the inference-framework PRs land, by Pavlichenko's account. Until then, any speed number labelled "+ MTP" is one you cannot reproduce.

## Why a 12B model runs like a small one

Two mechanisms carry the speed, and both were chosen during Mellum2's design, which the [technical report](https://arxiv.org/abs/2605.31268) describes as ablating every choice "under a fixed inference budget: matching the single-H100 speed of Qwen2.5-7B."

The first is routing. A [mixture-of-experts](/architectures/mixture-of-experts) layer replaces one wide MLP with many narrow ones and a router that picks a few per token. Here each expert is a SwiGLU MLP 896 wide: gate and up projections from 2,304 to 896, a down projection back. That is 6,193,152 parameters per expert, and 64 of them per layer. The router (`mlp.gate`, a 64-by-2,304 matrix) scores all 64, keeps the top 8, renormalises their weights (`norm_topk_prob: true`), and sums the 8 outputs. Decoding one token touches about a fifth of the weights. On a GPU that is memory-bandwidth bound during decoding, which is most of the time when you serve one user, the bytes you read per token set the speed. A 12B MoE that reads 2.4B parameters' worth of weights per token decodes like a 2.4B dense model, while holding 12B parameters' worth of knowledge.

The cost shows up somewhere else. All 12B parameters still have to sit in memory, and under heavy batching different tokens want different experts, so a saturated server ends up reading most of the experts anyway. The win under load comes from compute, not bandwidth: per token, the matrix multiplies are those of a 2.4B model.

The second mechanism is the cache. Each layer stores keys and values for past tokens: 4 KV heads, 128 wide, K and V, in bf16, which is 2,048 bytes per token per layer. In a sliding-window layer that cache stops at the last 1,024 tokens. In a full-attention layer it grows forever. With 21 sliding layers and 7 full ones, the sliding half caps out at 44 MB, and the per-token growth is 7 layers times 2,048 bytes, which is 14,336 bytes per token. At the full 131,072-token context that comes to 1.92 GB per sequence. Make all 28 layers full and it would be 7.52 GB, so the window buys a factor of 3.9 at long context.

<KvGrowth />

The comparison that matters is Qwen3.5-9B, because that is the model Mellum2.1 is chasing on the scorecard. Qwen3.5 caps three layers in four as well, using a Gated DeltaNet linear-attention layer with a fixed state instead of a sliding window. But its 8 full-attention layers use 256-wide heads, so they grow at 32,768 bytes per token, 2.3 times Mellum's rate. At 114,000 tokens, the context JetBrains used for its agentic evaluation, my arithmetic puts Mellum2.1 at 1.68 GB of cache per sequence and Qwen3.5-9B at about 3.79 GB. Smaller cache per sequence means more sequences fit in a batch, and more sequences in a batch is what "throughput under heavy load" measures.

The full-attention layers are also where the long context comes from, and the config shows a detail the report calls layer-selective YaRN. The `full_attention` layers carry YaRN scaling with factor 16 over an original 8,192 positions, which gets them to 131,072. The `sliding_attention` layers use plain RoPE with the same base of 500,000. A layer that never looks more than 1,024 tokens back has no reason to see rescaled positions, so only the layers that need to reach far are stretched.

The report's own measurements show what this bought for Mellum2. Against Qwen2.5-7B on one H100, it matched single-request speed to within one token per second (192 against 193) and served 21% more tokens per second under load (5,179 against 4,283).

<Figure
  src="https://ai.thesatyajit.com/articles/mellum2-1/fig4.png"
  alt="Two horizontal bar charts. Sync, single request: Qwen2.5-7B 193, Mellum 2 192, Qwen3-8B 169 output tokens per second. Throughput, concurrent: Mellum 2 5,179, Qwen2.5-7B 4,283, Qwen3-8B 2,897 output tokens per second."
  caption="Mellum2's speed target, met: equal to Qwen2.5-7B for one request and 21% ahead under load, on one H100 with vLLM FP8 serving at 2,304 input and 256 output tokens (Mellum2 technical report, Figure 13)."
/>

The report also plots why the window matters more as prompts grow. Its latency figure compares candidate MoE configurations with a 3:1 sliding-window pattern against Qwen2.5-7B at 2,304, 4,096 and 8,192 input tokens.

<Figure
  src="https://ai.thesatyajit.com/articles/mellum2-1/fig5.png"
  alt="Line chart of mean latency in seconds against input tokens at 2,304, 4,096 and 8,192. Qwen2.5-7B, dashed orange, rises from about 30 to about 72 seconds. Mellum 2 with hidden size 2304, 28 layers and window 1024, solid black, rises from about 22 to about 54 seconds. Speedup labels read 1.40x, 1.49x and 1.35x. Six other candidate configurations are drawn in grey near the Mellum 2 line."
  caption="Mean latency of the sliding-window MoE candidates against Qwen2.5-7B as the input grows; the chosen configuration is the black line (Mellum2 technical report, Figure 3)."
/>

None of this changed in 2.1. Which brings us to what did.

## Everything after pre-training

The blog's list of changes is short. Reinforcement learning "went from a short final stage to the main part of training." New RL tasks were added in math, competitive programming, science, tool use and software engineering, "combining open RL datasets with tasks we built ourselves," with every source filtered for broken tests, unverifiable answers, and tasks that are too easy or impossible. For software engineering, the model card adds, the model "trains inside real repositories with a shell and file-editing tools and is rewarded when the tests pass."

The more technical account is in Pavlichenko's thread, not the blog: "The team has scaled our RL infrastructure onto tens of millions of sandboxes, fixed training/inference mismatch issues, and replaced GRPO with CISPO." The blog says "millions of sandboxed runs across thousands of environments." Those two counts differ by an order of magnitude, and I have no way to check either. There is no technical report for 2.1 yet, so the dataset sizes, the number of RL steps, the clip thresholds and the environment list are all undisclosed.

What I can do is read the June report to see what was replaced. Mellum2's RL ran on NeMo-RL with Megatron-Bridge for training and vLLM for generation, asynchronous, with rollouts at most two policy versions stale. Rewards came from a separate verification gateway: a code sandbox for unit tests, a math verifier, an LLM judge for free-form answers. The loss was GRPO as most open labs now run it: token-level averaging, a leave-one-out baseline with no standard-deviation scaling, DAPO's clip-higher, no KL term. On top sat IcePop, which drops any token whose trainer-side probability disagrees too much with the probability the inference engine sampled it at.

That mismatch is the part Pavlichenko says they fixed, and the report explains where it comes from in an MoE: "for the same hidden state, the inference-time router may dispatch a token to a different expert than the trainer-side router." Two forward passes of the same weights pick different experts because of bf16 rounding near a routing boundary, and then the log-probabilities disagree. [Rollout Routing Replay](/articles/rollout-routing-replay) attacks the same problem by replaying the rollout's routing decisions in the trainer. JetBrains does not say which fix they used.

### Why CISPO

CISPO comes from the [MiniMax-M1 paper](https://arxiv.org/abs/2506.13585), and its motivation fits an agentic model well. PPO and GRPO clip the update: once a token's probability ratio $r = \pi_\theta / \pi_{\text{old}}$ leaves the band around 1, its gradient is cut to zero. MiniMax found that the tokens that leave the band first are the rare ones a model needs to learn: "tokens associated with reflective behaviors (e.g., However, Recheck, Wait, Aha), which often serve as 'forks' in reasoning paths." They start at low probability, so one update moves their ratio a lot, and then clipping silences them for the rest of the batch.

CISPO clips the importance weight instead and keeps every token's gradient:

$$
\mathcal{J}_{\text{CISPO}}(\theta) = \mathbb{E}\left[\frac{1}{\sum_i |o_i|}\sum_{i}\sum_{t} \operatorname{sg}\big(\hat r_{i,t}\big)\,\hat A_{i,t}\,\log \pi_\theta(o_{i,t}\mid q, o_{i,<t})\right],\quad \hat r_{i,t} = \operatorname{clip}\big(r_{i,t},\,1-\epsilon_{\text{low}},\,1+\epsilon_{\text{high}}\big)
$$

Here $\hat A_{i,t}$ is the group-relative advantage of response $i$ (the same one GRPO uses), $\operatorname{sg}$ is stop-gradient, and the sum runs over all tokens of all $G$ responses to one prompt. The ratio becomes a capped weight on a plain policy-gradient term. A token whose ratio ran away still pushes, it just pushes with a weight no larger than $1+\epsilon_{\text{high}}$. MiniMax set no lower bound in practice and tuned only the upper one.

For an agent trained on repositories, the forks are tokens like "let me run the tests first" or a decision to open a different file. They are rare in a base model and decisive for the outcome. A loss that keeps learning from them for the whole batch makes sense, and it is the same choice Poolside made for [Laguna](/articles/laguna-model-factory). Whether CISPO, the mismatch fix or simply ten times more environments produced the jump, JetBrains does not separate, and I would not guess.

One cost of the new recipe is already on record. Asked how fill-in-the-middle works with 2.1, Pavlichenko replied: "Please stick with Mellum2 base for FIM, I'd assume FIM is obliterated in this checkpoint." The tokenizer still carries `<fim_prefix>`, `<fim_middle>` and `<fim_suffix>`, but this checkpoint is a thinking agent now, not an autocomplete model.

## The scorecard

The model card's table has 17 benchmarks and four models, all run by JetBrains "with the same pipeline in thinking mode." The agentic rows are where the release is aimed.

| Agentic (Pi harness, 114K context) | Mellum2.1 | Mellum2 | Gemma 4 E4B | Qwen3.5-9B |
| :--- | ---: | ---: | ---: | ---: |
| SWE-bench Verified | 47.0 | 2.0 | 23.0 | 50.0 |
| Terminal-Bench 2.1 | 17.4 | 0.6 | 3.4 | 21.7 |
| SWE-bench Pro | 28.0 | 0.0 | 4.0 | 38.0 |

Mellum2 was, in Pavlichenko's words, "a good instruction model" that "struggled in complex agentic environments," and the table shows it: 2.0 on SWE-bench Verified is close to never solving an issue. Mellum2.1 solves 47 in 100 and roughly doubles Gemma 4 E4B. It still trails Qwen3.5-9B on all three, by 3.0 points on SWE-bench Verified, 4.3 on Terminal-Bench 2.1, and 10.0 on SWE-bench Pro, the hardest of the three. The summary chart picked SWE-bench Verified, where the gap is smallest.

<Figure
  src="https://ai.thesatyajit.com/articles/mellum2-1/fig3.png"
  alt="Bar chart titled SWE-bench Verified, tagged Agentic. Mellum2.1 Thinking 47.0, Mellum2 Thinking 2.0, Gemma 4 E4B Thinking 23.0, Qwen3.5-9B Thinking 50.0, on an axis from 0 to 60 percent."
  caption="SWE-bench Verified in JetBrains' own harness: Mellum2.1 at 47.0, three points under Qwen3.5-9B (JetBrains blog, SWE-bench Verified card)."
/>

The harness is worth naming. The card says all agentic runs use "the same open-source agent harness (Pi v0.73.1, shell and file tools) for every model, with each model's default sampling (temperature 1.0 for Mellum2.1), a 114K-token context, and up to 16K tokens per turn." Pi is Mario Zechner's minimal coding agent; version 0.73.1 is on npm as `@mariozechner/pi-coding-agent`, published 7 May 2026, with read, bash, edit and write tools. A [harness](/articles/agent-harness) with four tools and no repository index is a fair floor for comparing small models, but scores from it are not comparable with SWE-bench numbers that labs report from their own scaffolds. Read the 47.0 against the other three columns in this table, not against a leaderboard.

Outside the agentic rows, 2.1 improves on Mellum2 almost everywhere: LiveCodeBench v6 from 69.4 to 82.0 (the best of the four), AIME 25/26 from 60.1 to 83.3, BFCL v4 from 49.6 to 62.3, IFEval from 79.5 to 90.6. HarmBench's harmful rate falls from 21.5 to 8.5. Two rows go the other way, slightly: WorkBench from 45.1 to 44.6, and XSTest (answering safe prompts that look unsafe) from 91.2 to 88.8. Knowledge is still the weak spot it was in June: on GPQA Diamond, 64.6 against Qwen3.5-9B's 77.8.

### The baseline moved

One thing I noticed only by putting the June report next to the new card. The card notes that Mellum2 Thinking "was re-evaluated with this pipeline, so its numbers differ slightly from the Mellum2 Technical Report." Some differ more than slightly. The report's Table 10 gave the RL'd Mellum2 Thinking 57.6 on GPQA Diamond; the new card gives it 51.0. So the "+13.6 vs Mellum2" printed under GPQA in the summary chart would be +7.0 against the June number.

Qwen3.5-9B moved further. In June's table it scored 73.4 on AIME, 42.7 on BFCL v4 and 68.3 on LiveCodeBench v6. In October's it scores 86.7, 58.5 and 75.4. Same model, same company running it, a 13.3-point swing on AIME. The June report explains one likely cause in its appendix: small Qwen thinking models often never close their `</think>` tag, so JetBrains imposed a 32K-token reasoning budget. The new card says only that non-agentic benchmarks use greedy decoding. I can't tell what changed in the pipeline. I take one thing from it: the gaps in any single table are only as stable as the harness behind them, and here the moves went both ways, Mellum2 down and Qwen3.5-9B up, so I don't read the new table as tilted.

## Which speed

<Figure
  src="https://ai.thesatyajit.com/articles/mellum2-1/fig2.png"
  alt="Chart titled Model comparison, output tokens per second, FP8, input and output lengths 2304 and 256, one H200. Left half, Sync, one request at a time: Mellum2.1 plus MTP 557, Mellum2.1 339, Qwen3.5-9B plus MTP 426, Qwen3.5-9B 241, Gemma 4 E4B plus drafter 607, Gemma 4 E4B 232. Right half, Throughput, server saturated: 8,533, 7,969, 4,373, 4,347, 6,618 and 6,099 respectively."
  caption="JetBrains' speed comparison: output tokens per second on one H200 in FP8, at 2,304 input and 256 output tokens, for one request at a time and with the server saturated (JetBrains blog, Model comparison chart)."
/>

This is the chart "relative to its speed" points at, and the conditions are printed across the top: output tokens per second, FP8, input and output lengths of 2,304 and 256, one H200. Those lengths are the ones the Mellum2 report calls "representative input/output sizes from production code completion workloads." It is the benchmark rig from June, moved to newer hardware. Because the GPU changed (H100 then, H200 now), Mellum2's 192 tokens per second and Mellum2.1's 339 are not comparable, and the blog does not compare them.

Under load, the blog's claims check out against its own bars. Mellum2.1 serves 7,969 tokens per second against Qwen3.5-9B's 4,347, a ratio of 1.83, which is "almost twice." With drafts on, 8,533 against 4,373. For a single request, MTP takes Mellum2.1 from 339 to 557, a factor of 1.64, which the blog rounds to "about 1.6 times faster."

The chart shows two more things the prose leaves out. For a single request with drafts on, the fastest model in the group is Gemma 4 E4B with its drafter, at 607. And MTP barely helps any model under saturation, adding 7% to Mellum2.1 and under 1% to Qwen3.5-9B. That is expected: speculative decoding spends spare compute to cut latency, and a saturated server has no spare compute.

Put the scores and the speeds on one plot and "relative to its speed" becomes a precise statement. A model is on the frontier if no other model is both faster and better. Pick a benchmark and a speed mode:

<ScoreSpeedFrontier />

On SWE-bench Verified with the server saturated, Mellum2.1 and Qwen3.5-9B are both on the frontier and Gemma 4 E4B is not: Mellum2.1 is both faster and better than it. Mellum2.1 is the fastest point but not the best one. Switch to SWE-bench Pro and the picture holds, with a ten-point gap instead of three. On LiveCodeBench v6 or BFCL v4, Mellum2.1 is both fastest and best, and stands alone on the frontier. On WorkBench, with one request at a time, Mellum2 at the same speed edges out its successor, which is what a 0.5-point regression looks like on this plot. The claim, read carefully, is that Mellum2.1 is never dominated on an agentic benchmark. That is true in JetBrains' numbers, and it is a weaker claim than the headline.

What the chart cannot tell you is how the speeds look at the context these agents actually run at. The agentic evaluations used a 114K-token context with up to 16K tokens per turn. The speed test used 2,304 tokens in. At long context, attention and cache size matter more and weight reads matter less, and that is where Mellum2.1's 14,336 bytes per token against Qwen3.5-9B's 32,768 should count. My guess is that the throughput gap would widen at agent-length contexts, because more Mellum sequences fit in the same memory. That is a guess from the config arithmetic. JetBrains did not measure it, and neither did I.

## Running it

vLLM serves it directly. The card's command uses the Qwen3 reasoning parser and the Hermes tool-call parser, which matches the chat template: `<think>` blocks and `<tool_call>` JSON inside ChatML turns.

```sh
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes
```

The card recommends temperature 0.6, top-p 0.95 and top-k 20 for chat, but the agentic results were run at temperature 1.0, so if you are reproducing the SWE-bench number, use that.

Both the blog and the main model card say GGUF builds are "coming soon." They are already up: the GGUF repository was created on 7 October and holds five files. llama.cpp has a `mellum` architecture in `src/llama-arch.cpp`, which is the string the GGUF header carries. The card measures each quant against bf16 logits on Wikitext-2:

| File | Size | KLD vs BF16 | Top-token match |
| :--- | ---: | ---: | ---: |
| BF16 | 24.3 GB | n/a | n/a |
| Q8_0 | 12.9 GB | 0.008 | 96.1% |
| Q6_K | 10.9 GB | 0.020 | 93.9% |
| Q4_K_M (recommended) | 8.1 GB | 0.075 | 88.0% |
| MXFP4_MOE | 7.0 GB | 0.115 | 85.6% |

I read the tensor tables out of the Q4_K_M and MXFP4_MOE headers, and the Q4_K_M file is heavier than its name suggests: 5.31 bits per parameter rather than the 4.5 or so you would expect. The reason is a shape. Every expert's down projection has rows 896 wide, and llama.cpp's k-quants work in blocks of 256, which 896 does not divide. So the down projections fall back to non-k formats: Q8_0 in 14 layers and Q5_0 in the other 14, while the gate and up projections get Q4_K. The Mellum2 report says JetBrains kept every dimension "divisible by 128 or higher powers of two" for GPU kernels. 896 is 7 times 128, which is fine for a GPU kernel but one factor of two short of what a k-quant needs. MXFP4_MOE quantizes every expert to MXFP4 (blocks of 32, so 896 is fine) and keeps attention and embeddings at Q8_0, for 4.63 bits per parameter and 7.03 GB.

```sh
llama-server -hf JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M \
  --ctx-size 131072 --temp 0.6 --top-p 0.95 --top-k 20
```

An 8 GB file plus a 1.9 GB cache at full context fits a 12 GB card. On a 16 GB Mac it fits too, and because only 2.44B parameters are read per token, decoding stays quick on hardware where a dense 9B model would be slow. I have not timed it; the speeds in this article are JetBrains'.

## Where I'd use it

As a sub-agent. The blog's own list of uses starts with "a capable worker inside agentic systems," and the numbers back that framing more than they back "a coding agent" on its own. A planner on a frontier model that hands Mellum2.1 a bounded task, like finding the root cause of a failing test or drafting a fix for review, gets nearly Qwen3.5-9B's agentic quality at almost twice its serving throughput, under Apache 2.0, on one GPU. For a fleet of parallel workers on your own hardware, that trade is the point of the model.

As the only agent on a hard repository, I'd still pick something larger, and on SWE-bench Pro in this very table Qwen3.5-9B is ten points ahead. For completion inside an editor, use Mellum2 base, as its author says.

I'm watching for three things. The MTP head, because the single-request numbers depend on it. A speed test at agent-length context, which I expect to favour Mellum2.1 more than the published one does. And the 2.1 technical report, because "tens of millions of sandboxes" and a GRPO-to-CISPO switch on a 2.5B-active MoE is a recipe other small-model teams would want to copy, and right now it is two sentences in a tweet.

## How I checked

I read Pavlichenko's thread and its replies through the fxtwitter mirror, the JetBrains blog post, both Hugging Face model cards, `config.json`, `chat_template.jinja` and `tokenizer_config.json`. Parameter counts come from reading all five safetensors headers by HTTP range request and summing tensor shapes; the total matches the index file's `total_parameters`. The active count is non-expert parameters plus 8/64 of expert parameters. The architecture diff is a key-sorted comparison of this `config.json` with Mellum2 Thinking's. GGUF details come from parsing the metadata and tensor tables of the Q4_K_M and MXFP4_MOE files, fetched with a range request over their first 12 MB. File sizes come from the Hub API. KV-cache figures are arithmetic on the two configs (bf16 cache, one sequence; Qwen3.5-9B's linear-attention state counted in fp32 from its `mamba_ssm_dtype`, its small convolution state left out). The June baselines come from Table 10 and Figure 13 of the Mellum2 technical report, and the CISPO objective and quote from Section 3 of the MiniMax-M1 paper. The Pi version and date come from the npm registry. I did not run the model; every benchmark score and tokens-per-second figure is JetBrains', and the frontier plot only rearranges them.
