2026-10-06 · 18 min · calibration · benchmarks · agents
Why read this
Hightop 30%Byte diffs show Agr is Gemma plus a merged adapter; the control the launch chart omits puts the fine-tune at +0.82, almost none of it in Tools.
- Original, source-checked analysis
- Concrete numbers to act on
- Open code or weights
LLM architectureNeeds a workstation GPUApache-2.0Practitioner model
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 3 of 3: Reproduces the headline result, or shows from primary files it is wrong
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 3 of 3: The only place this analysis exists
Score 69 of 100, ranked 112 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
Command Code's launch post is four lines and a chart: "Agr (31B) and Agr-flash (360M)", "58.15 on Decision Index 0.2.1", "Focus on: tool calls & routing", "TypeSafe SDK support". Then the one sentence that makes it a decision model: "Agr skips text generation and returns typed values with a probability for each option." The thread adds that the training runs focused "on accuracy for tool calls and coding-task routing", and links the Hugging Face page.
The category is not new to this site. Week four covered the readout families and the softmax-and-temperature step, and the llama.cpp piece traced a request from JSON to number and showed that "one forward pass" there means one pass per question. I will not repeat either. This piece is about what is specific to Agr: which weights it actually ships, how one pass answers every question without the questions seeing each other, where the probabilities come from and how they were tuned, and whether 58.15 means what the chart implies.
Everything below is labelled. Measured means I computed it from a file: safetensors headers and tensor bytes read with HTTP range requests, the configs, the code at commit 296303d, and the Decision Index's own data/index.json. Reported means Command Code's figure, not re-run. Reasoned is my arithmetic on the other two. I did not run either model; the 31B needs a 96 GB GPU by its own README, and my sandbox has no PyTorch.
- task
- text-classification
- library
- agr
- license
- apache-2.0
- safetensors
- 13 shards
- largest file
- 5.00 GB
- files
- 26
- downloads
- 82
- likes
- 7
Gemma 4 31B IT, text tower only: 30,697,345,340 parameters in bf16 across 13 shards (measured from the headers). Linear projections fine-tuned and merged; embeddings, norms and layer scalars bit-identical to the base. Readout: letters, with a fitted temperature in config.json. Apache-2.0.
repo last modified 2026-10-05
- task
- text-classification
- library
- agr
- license
- apache-2.0
- safetensors
- 2 shards
- largest file
- 723.7 MB
- files
- 13
- downloads
- 97
- likes
- 1
SmolLM2-360M backbone (361,821,120 parameters, retrained) plus head.safetensors: five marker embeddings and 960-to-256 query and key projections, 496,320 parameters in fp32 (measured). No published benchmark score. Apache-2.0.
repo last modified 2026-10-05
- license
- Apache-2.0
- branch
- main
- tests
- 4 files
- source
- 85.4 kB
- commit date
- 2026-10-05
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
Read at 296303d (5 October 2026): agr/model.py (301 lines) is the whole readout; DESIGN.md is the design note. Not on PyPI: install from git as commandcode-agr.
local clone, 2026-10-06 at 296303d — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
What is in the two repositories
The model cards name the bases in their metadata: google/gemma-4-31B-it for Agr and HuggingFaceTB/SmolLM2-360M for Agr-flash. Both bases are Apache 2.0, and so are both Agr checkpoints (measured, from the Hub tags and LICENSE.md). The NOTICE.md says what changed: "fine-tuned weights merged into the base model", with Agr keeping "the base model's vocabulary, chat template and output layer".
I recounted the parameters from the safetensors headers rather than trusting the names (measured):
| Agr | Agr-flash | |
|---|---|---|
| Backbone | Gemma 4 31B IT, text tower | SmolLM2-360M |
| Backbone parameters | 30,697,345,340 | 361,821,120 |
| Extra head | none | 496,320 (fp32) |
| Layers, hidden size | 60, 5,376 | 32, 960 |
| Vocabulary | 262,144 | 49,152 |
| Weight bytes on the Hub | 61,427,769,000 | 729,910,317 |
| Token limits (state / question / request) | 16,384 / 16,384 / 32,768 | 6,144 / 2,048 / 12,288 |
So "31B" is 30.70B, and "360M" is 362.3M with the head (reasoned). Both are the honest rounding. The shard index's own metadata says 30,697,345,280, which is 60 fewer than the headers sum to: exactly the 60 one-element layer_scalar tensors, one per layer (measured). The README's 61.43 GB and 0.73 GB match the repositories' byte totals.
The Gemma checkpoint Agr started from is larger. google/gemma-4-31B-it is 32,682,372,656 parameters, 355 of its tensors in a vision tower (measured, from its index). Agr ships only the 832 language-model tensors. That matches the card's "English text only; images are not supported", and it is why Agr is not a 32.7B model.
What the fine-tune touched
A merged fine-tune leaves no adapter file to inspect, so I compared tensors directly: the first 512 KiB of a tensor from each repository, fetched by byte range, decoded from bf16 (measured).
embed_tokens, the finalnorm, everyinput_layernormand everylayer_scalarI read are bit-identical to Gemma's.- Every linear projection I read differs: q, k, v and o in attention, gate, up and down in the MLP, at layers 0, 5, 30 and 59. The relative change, ‖Agr − Gemma‖ / ‖Gemma‖ over the slice, is 0.12% to 0.74%, and 30% to 75% of the individual bf16 values are unchanged.
That is the fingerprint of a low-rank adapter trained on the projections and merged back in (reasoned; the card says only "fine-tuned weights merged"). It matters for the readout: Gemma ties its output layer to its input embeddings, so an untouched embedding table means Agr reads its answers through exactly Gemma's output layer. Everything Agr learned lives in the hidden state at the answer position.
Agr-flash is a different story. Its projections differ from SmolLM2-360M's by 4% to 15% and almost no values survive (measured, layers 0, 16 and 31), while the embeddings and norms are again identical. That is a much heavier retraining of a much smaller model. The NOTICE.md says Agr-flash "adds marker tokens to the vocabulary", but the embedding table is still 49,152 rows (measured); the five markers are vectors in head.safetensors that the code substitutes into the input embeddings at their positions (_inputs() in agr/model.py).
Scoring instead of generating
A chat model asked "which team should handle this?" writes an answer token by token, and you parse it. Its probability of "billing" is spread over every way of spelling, prefixing or explaining billing, and over every token that is not an answer at all.
A decision model removes the choice of what to write. Agr's letters readout puts the options into the prompt under single-token codes and ends the prompt where the model's reply would begin:
<user turn> Context:
<the state>
Question: Which team should handle this?
Options:
(A) Engineering: technical faults
(B) Billing: payment disputes <end of turn> <model turn> Answer: (Then it runs the backbone once and takes the hidden vector at the last token, the (. Gemma's next-token logits there are , one row per vocabulary entry. Agr keeps only the rows of the option codes :
where is the embedding row of code , 30 is Gemma's final_logit_softcapping, and is the fitted temperature for options (measured, _letters() in agr/model.py). That is the whole readout: a softmax over entries of the vocabulary, renormalised so they sum to one. Nothing is sampled, nothing is parsed, and the output cannot be anything but one of the options.
Codes are A to Z, then two-letter pairs, kept only when the tokenizer makes them one token, up to 255 (measured, option_codes()). Above ten options each code is spliced in as its token id, so a pair like AB cannot be split by its neighbours. A noul is a two-option choice whose options read "no" and "yes"; its answer is . A score lists its levels as 0: …, 1: … and returns the expected level .
This is the Decider layout, and the card says so. The NOTICE.md credits Decider by Mark Marosi for "the answer-slot prompt layout, the single-token option codes … and the temperature as a function of the number of options". Remember that credit; it matters for the benchmark.
Agr-flash has no chat format to borrow, so it uses a trained head. Its layout wraps the state in a <doc> marker and each option in <choice>…</choice>, and ends each question with <answer>. The hidden state at <answer> becomes a query, each option's tokens are averaged into a key, and the logits are scaled dot products (measured, Scorer.forward()):
and are 256 by 960. Two of them plus five 960-wide markers is the 496,320-parameter head (reasoned: ). It is the same pointer idea as Kev's, except that here each option is pooled from its own span rather than read at one token, and Agr-flash has no fitted temperature in its config at all.
One pass for every question
The llama.cpp endpoint answers a three-question request with three prompts; with a free slot it copies the state's cache, and without one it reads the state again. Agr puts the whole request in one sequence and changes the attention mask so that it behaves like separate prompts (measured, Layout.visibility()):
- The state is ordinary causal attention.
- Each question's block sees the whole state and its own earlier tokens, and never another question's block.
- Each block's position ids restart where the state ends, "as if it were the only question".
So the state is read once, every question is a branch hanging off it, and the answer is read at the last token of each branch. Gemma's sliding-window layers (50 of the 60, window 1,024, measured from layer_types) get the same mask, further limited by position distance; the code builds a separate mask for each layer type and relies on SDPA to apply it. Toggle the plain causal layout to see what the mask removes, and change the sizes to see what reading the state once saves:
row = token reading · column = token read · top numbers = position ids
Every question block reads the state and itself. Its position ids restart at 6, as if it were the only question, so adding, removing or reordering questions leaves each answer alone.
The cost difference is plain arithmetic. With a state of tokens and questions of tokens each, one pass reads tokens; separate prompts with nothing shared read . For a 2,000-token state and three 60-token questions that is 2,180 against 6,180 (reasoned). The design note's claim is that the answers do not change either way: in fp32 an answer is the same whichever other questions come with it "to within 0.00001", and bf16 on a GPU moves answers "by up to about 0.02" (reported; I did not run it). The design is credited to TypeSafe's Jev as described in Archer Hume's "Jev's Architecture Unmasked", and to Kev.
What the mask does not remove is option order. Within one question the options are still read in sequence, and A and B still carry different priors. The code can read each question twice with the options reversed and average the two (order_averaging, credited to Invergent AI's Surogate), but Agr's config.json does not set it (measured), and the card's limitations say it outright: "Changing the order of options can change the answer."
A server caps a request at 64 questions and 1 MB of body; a state or question over its token budget is cut, and the response says "truncated": true; a request over the total is refused with a 422 (measured, agr/serve.py).
Where the probabilities come from, and how they were tuned
"A probability for each option" is a claim about calibration: that 0.8 means right about 80% of the time. Agr's only calibration mechanism is the temperature. Its config.json stores
"temperature_by_options": { "a": 1.25, "b": -0.05, "min": 0.05 }and the code computes , "fitted by log loss on our development data" (reported, DESIGN.md). That gives 1.215 for a yes/no question, 1.181 for four options and 0.973 for 255 (reasoned). A temperature above 1 flattens the distribution; Agr's raw logits are slightly overconfident for small .
The comparison that makes this interesting is Decider's own checkpoint of the same model. Mapika/decider-chat-gemma4-31b is "stock Gemma-4-31B-it" read through the same layout with "unmodified weights" and its own fitted curve (measured, its decider_config.json):
"temperature_by_options": { "a": 10.124, "b": -1.633, "min": 0.05 }That is 8.992 at two options, 7.86 at four and 4.804 at twenty-six (reasoned). Stock Gemma's letter logits are so sharp that a yes/no answer has to be divided by nearly nine before its probabilities mean anything. After Agr's fine-tune, the same output layer needs about 1.2. The fine-tune did not just change which answer wins; it trained the hidden state toward logits that are already near the right scale, which is what training on log loss over the option codes does (reasoned; the training recipe is not published).
x: options (log) · y: temperature · Agr · decider-chat, stock weights
top option ahead of 3 equal rivals by the gap (illustrative)
Two cautions. First, a single curve per option count corrects average overconfidence; it cannot fix a model that is overconfident on one kind of question and underconfident on another. Second, Agr publishes no calibration number. The Decision Index measures one for every board entrant (a 1-in-6 sample of 32 benchmarks, confidence as the probability on the chosen option): Jev's expected calibration error is 0.074, decider-chat on stock Gemma 0.047, and pplx-decider-v1-27b 0.0178 (measured, from data/index.json). Agr is not on the board, so it has no such figure. The card's own advice is the right one: check thresholds on your own labelled data, which agr eval reports as accuracy, Brier score and calibration error.
The confidence field in a response is not a calibrated probability either. For a choice it is System One's (p_\max - 1/K) / (1 - 1/K): 0 for a uniform spread, 1 for certainty. On four options a top probability of 0.7 is a confidence of 0.6 (reasoned).
Checking the 58.15

The headline holds as arithmetic. The Decision Index 0.2.1 is a weighted mean of five areas' chance-corrected skill: Knowledge 25.8%, Language 25.8%, Retrieval 20.0%, Tools 18.3%, Arts 10% (measured, index02.py in the kit at 87d4650). Applying those weights to the board's area scores for Jev gives 57.91, its board figure. Applying them to Agr's five rounded area scores from its own chart gives 58.14, consistent with 58.15 after rounding (reasoned). One reply asked whether 58.1 on the chart and 58.15 in the post disagree; they do not, the chart rounds to one place. The "150,759 requests" on the card is the kit's 120,340 rescored rows plus the 30,419 added in 0.2 (measured, the kit's README and tests).
What the number is not is a board entry. The footer says "Agr: our own run of the public kit". The board's index.json was generated on 28 September and has 70 entrants and no Agr (measured). A reply asked whether Agr will be scored on the public board; nothing says so yet. A quarter of a point over Jev (58.15 against 57.91, reasoned) from a self-run is a tie until someone else runs it.

Per area, Agr is ahead of Jev on Retrieval (65.3 against 55.4), Tools (75.9 against 75.1) and Arts (39.7 against 37.7), and behind on Knowledge (45.8 against 51.4) and Language (59.5 against 62.0) (reported for Agr, measured for Jev). The Knowledge gap is reasoning: GPQA Diamond 34.0 against 71.4, BBH 66.5 against 89.7 (same labels). Across the 38 benchmarks Agr is ahead of Jev on 21 and behind on 17 (reasoned, from the per-benchmark chart below). The Jev values in Command Code's charts match the board file to the decimal I checked (measured).
The comparison the chart leaves out
The board's second place, at 57.33, is "Decider chat · Gemma-4-31B": stock google/gemma-4-31B-it, no training, read through Decider's chat layout with the 10.124 temperature curve (measured, data/index.json and the decider-chat config). That is Agr's backbone and Agr's readout without Agr's fine-tune. It is the control experiment, it was on the board Command Code copied its other bars from, and the chart shows Jev, Rune and Kev instead.
Put Agr's area scores next to it (reported for Agr, measured for decider-chat, differences reasoned):
| Area | Weight | Agr | Same Gemma, no training | Change |
|---|---|---|---|---|
| Knowledge | 25.8% | 45.8 | 44.3 | +1.5 |
| Language | 25.8% | 59.5 | 60.4 | −0.9 |
| Retrieval | 20.0% | 65.3 | 63.1 | +2.2 |
| Tools | 18.3% | 75.9 | 75.6 | +0.3 |
| Arts | 10% | 39.7 | 38.3 | +1.4 |
| Index | 58.15 | 57.33 | +0.82 |
The fine-tune is worth 0.82 index points over the same weights read the same way. That is real; 26 of 38 benchmarks go up, 10 go down and 2 do not move (reasoned). But it is not where the launch post says the effort went. The Tools area, five benchmarks, moves 0.3: API-Bank 84.8 to 85.0, BFCL 97.3 to 97.3, Home appliances 62.5 to 63.6, ToolRet 58.0 to 60.3, When2Call 69.2 to 67.2. The big gains are elsewhere: Amazon ESCI and HoVer up 7.3 each, SATA-Bench 7.2, ForecastBench 9.1. The big losses are NLI4CT down 9.5, PhishNChips down 7.5 and iSarcasmEval down 7.1.

"Best-in-class on coding agent tool calls and routing" needs two corrections (reasoned, from the board). The Decision Index has no coding-agent benchmark, so that part cannot be checked against it at all; the Tools area is general function calling and API use. And on that area, three board entrants already score higher than Agr's 75.9: pplx-decider-v1-27b at 79.3, Jebadiah 27B at 78.1 and simple-jev on Qwen3.8-27B at 76.2 (measured). Where Agr does lead every board entrant is the area its chart calls "Retrieval and routing": 65.3 against Rune's 63.5. If "routing" means that, the claim holds there.
Agr-flash, unscored
Several replies asked for Agr-flash's number. There is none. The card gives it a row in the models table, token limits and the line "Agr-flash is not meant for safety decisions", and every chart is Agr. A 362M model with a trained head is the interesting deployment shape, the one the replies kept pointing at, and nothing published says how much it gives up. For scale, the 0.8B Kev on the same board scores 14.6 (measured); I would not guess where Agr-flash lands.
Smaller things the files say
- TypeSafe SDK support means Agr's server speaks System One's request and response format at
/v1/systemone, so TypeSafe's own SDKs (typesafe-sdkon PyPI,@typesafe-ai/sdkon npm, both TypeSafe AI's) work against it withbase_urlpointed at localhost (measured, the package registries and the card). It is a compatible wire format, not a Command Code SDK. Command Code's code is the Python package in CommandCodeAI/agr, installed from git ascommandcode-agr;agron PyPI is an unrelated project, aspyproject.tomlnotes. - The version pin is off by one.
pyproject.tomlpinstransformers>=5.17,<5.18because Agr "relies on per-layer-type attention masks … which are not stable API", and both backbones'config.jsonfiles say they were written by5.18.0(measured). Probably harmless; worth knowing if a load fails. - Requests are serialized per GPU: "Requests on the same GPU run one at a time" (reported). On a 96 GB card, that is one 61 GB model answering one request at a time. The card publishes no latency; the board measures decider-chat on stock Gemma at a 108.5 ms median, and Jev's hosted API at 524.1 ms (measured,
data/index.json), but neither is Agr.
What holds
The mechanism is clean and documented: one pass over the state, a masked branch per question, Gemma's own output layer restricted to the option codes, a fitted temperature, and a design note that says which pieces came from Decider, Jev, Kev and Surogate. Both checkpoints are what the card says, at the sizes it says, under Apache 2.0, and the fine-tune is visible in the bytes. The temperature comparison is the clearest evidence of what training a decision model does to a chat model: the same output layer goes from needing a divisor of about nine to needing about 1.2.
The 58.15 is a self-run quarter-point over Jev, and the chart omits the one bar that shows how much of it is the fine-tune: 0.82 points over the same Gemma read the same way, with almost none of it in Tools. If you are choosing a decision model for tool routing, the Decision Index's Tools column says Agr is in the top group and not at the top of it, and the card's last advice applies to every model in this category: measure on your own labelled requests before trusting a threshold.
Sources, read on 6 October 2026: Command Code's post, its thread and first page of replies through the fxtwitter mirror; CommandCode/agr and CommandCode/agr-flash (cards, configs, NOTICE.md, safetensors headers and byte ranges); CommandCodeAI/agr at 296303d (agr/model.py, agr/serve.py, DESIGN.md, pyproject.toml); google/gemma-4-31B-it and HuggingFaceTB/SmolLM2-360M byte ranges; Mapika/decider-chat-gemma4-31b decider_config.json and Mapika/decider; the Decision Index kit at 87d4650 and the board's data/index.json (generated 28 September 2026) from the Decision Index space.