~/satyajit

Agr: a decision model is Gemma's own output layer, read at one position per question

mdjsonmcp

2026-10-06 · 18 min · calibration · benchmarks · agents

Why read this

Hightop 30%

Byte diffs show Agr is Gemma plus a merged adapter; the control the launch chart omits puts the fine-tune at +0.82, almost none of it in Tools.

  • Original, source-checked analysis
  • Concrete numbers to act on
  • Open code or weights

LLM architectureNeeds a workstation GPUApache-2.0Practitioner model

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
3 of 3: Reproduces the headline result, or shows from primary files it is wrong
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
1 of 3: A specialist community
Only here?
3 of 3: The only place this analysis exists

Score 69 of 100, ranked 112 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Command Code's launch post is four lines and a chart: "Agr (31B) and Agr-flash (360M)", "58.15 on Decision Index 0.2.1", "Focus on: tool calls & routing", "TypeSafe SDK support". Then the one sentence that makes it a decision model: "Agr skips text generation and returns typed values with a probability for each option." The thread adds that the training runs focused "on accuracy for tool calls and coding-task routing", and links the Hugging Face page.

The category is not new to this site. Week four covered the readout families and the softmax-and-temperature step, and the llama.cpp piece traced a request from JSON to number and showed that "one forward pass" there means one pass per question. I will not repeat either. This piece is about what is specific to Agr: which weights it actually ships, how one pass answers every question without the questions seeing each other, where the probabilities come from and how they were tuned, and whether 58.15 means what the chart implies.

Everything below is labelled. Measured means I computed it from a file: safetensors headers and tensor bytes read with HTTP range requests, the configs, the code at commit 296303d, and the Decision Index's own data/index.json. Reported means Command Code's figure, not re-run. Reasoned is my arithmetic on the other two. I did not run either model; the 31B needs a 96 GB GPU by its own README, and my sandbox has no PyTorch.

CommandCode/agr@cc47f0b · snapshot 2026-10-06
repo size
61.43 GB
task
text-classification
library
agr
license
apache-2.0
safetensors
13 shards
largest file
5.00 GB
files
26
downloads
82
likes
7
agrdecision-modelclassifier

Gemma 4 31B IT, text tower only: 30,697,345,340 parameters in bf16 across 13 shards (measured from the headers). Linear projections fine-tuned and merged; embeddings, norms and layer scalars bit-identical to the base. Readout: letters, with a fitted temperature in config.json. Apache-2.0.

repo last modified 2026-10-05

CommandCode/agr-flash@ca7b6a2 · snapshot 2026-10-06
repo size
726.7 MB
task
text-classification
library
agr
license
apache-2.0
safetensors
2 shards
largest file
723.7 MB
files
13
downloads
97
likes
1
agrdecision-modelclassifier

SmolLM2-360M backbone (361,821,120 parameters, retrained) plus head.safetensors: five marker embeddings and 960-to-256 query and key projections, 496,320 parameters in fp32 (measured). No published benchmark score. Apache-2.0.

repo last modified 2026-10-05

CommandCodeAI/agr@296303d · snapshot 2026-10-06
tracked files
24
license
Apache-2.0
branch
main
tests
4 files
source
85.4 kB
commit date
2026-10-05
source by language
Python64.6 kB(12)HTML20.8 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

Read at 296303d (5 October 2026): agr/model.py (301 lines) is the whole readout; DESIGN.md is the design note. Not on PyPI: install from git as commandcode-agr.

local clone, 2026-10-06 at 296303d — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

What is in the two repositories

The model cards name the bases in their metadata: google/gemma-4-31B-it for Agr and HuggingFaceTB/SmolLM2-360M for Agr-flash. Both bases are Apache 2.0, and so are both Agr checkpoints (measured, from the Hub tags and LICENSE.md). The NOTICE.md says what changed: "fine-tuned weights merged into the base model", with Agr keeping "the base model's vocabulary, chat template and output layer".

I recounted the parameters from the safetensors headers rather than trusting the names (measured):

AgrAgr-flash
BackboneGemma 4 31B IT, text towerSmolLM2-360M
Backbone parameters30,697,345,340361,821,120
Extra headnone496,320 (fp32)
Layers, hidden size60, 5,37632, 960
Vocabulary262,14449,152
Weight bytes on the Hub61,427,769,000729,910,317
Token limits (state / question / request)16,384 / 16,384 / 32,7686,144 / 2,048 / 12,288

So "31B" is 30.70B, and "360M" is 362.3M with the head (reasoned). Both are the honest rounding. The shard index's own metadata says 30,697,345,280, which is 60 fewer than the headers sum to: exactly the 60 one-element layer_scalar tensors, one per layer (measured). The README's 61.43 GB and 0.73 GB match the repositories' byte totals.

The Gemma checkpoint Agr started from is larger. google/gemma-4-31B-it is 32,682,372,656 parameters, 355 of its tensors in a vision tower (measured, from its index). Agr ships only the 832 language-model tensors. That matches the card's "English text only; images are not supported", and it is why Agr is not a 32.7B model.

What the fine-tune touched

A merged fine-tune leaves no adapter file to inspect, so I compared tensors directly: the first 512 KiB of a tensor from each repository, fetched by byte range, decoded from bf16 (measured).

That is the fingerprint of a low-rank adapter trained on the projections and merged back in (reasoned; the card says only "fine-tuned weights merged"). It matters for the readout: Gemma ties its output layer to its input embeddings, so an untouched embedding table means Agr reads its answers through exactly Gemma's output layer. Everything Agr learned lives in the hidden state at the answer position.

Agr-flash is a different story. Its projections differ from SmolLM2-360M's by 4% to 15% and almost no values survive (measured, layers 0, 16 and 31), while the embeddings and norms are again identical. That is a much heavier retraining of a much smaller model. The NOTICE.md says Agr-flash "adds marker tokens to the vocabulary", but the embedding table is still 49,152 rows (measured); the five markers are vectors in head.safetensors that the code substitutes into the input embeddings at their positions (_inputs() in agr/model.py).

Scoring instead of generating

A chat model asked "which team should handle this?" writes an answer token by token, and you parse it. Its probability of "billing" is spread over every way of spelling, prefixing or explaining billing, and over every token that is not an answer at all.

A decision model removes the choice of what to write. Agr's letters readout puts the options into the prompt under single-token codes and ends the prompt where the model's reply would begin:

<user turn> Context:
<the state>
 
Question: Which team should handle this?
Options:
(A) Engineering: technical faults
(B) Billing: payment disputes <end of turn> <model turn> Answer: (

Then it runs the backbone once and takes the hidden vector hh at the last token, the (. Gemma's next-token logits there are WEhW_E h, one row per vocabulary entry. Agr keeps only the rows of the option codes c1,…,cKc_1, \dots, c_K:

zk=30tanh⁡ ⁣(wck⊤h30),pk=exp⁡(zk/T(K))∑j=1Kexp⁡(zj/T(K))z_k = 30 \tanh\!\left(\frac{w_{c_k}^\top h}{30}\right), \qquad p_k = \frac{\exp(z_k / T(K))}{\sum_{j=1}^{K} \exp(z_j / T(K))}

where wckw_{c_k} is the embedding row of code ckc_k, 30 is Gemma's final_logit_softcapping, and T(K)T(K) is the fitted temperature for KK options (measured, _letters() in agr/model.py). That is the whole readout: a softmax over KK entries of the vocabulary, renormalised so they sum to one. Nothing is sampled, nothing is parsed, and the output cannot be anything but one of the KK options.

Codes are A to Z, then two-letter pairs, kept only when the tokenizer makes them one token, up to 255 (measured, option_codes()). Above ten options each code is spliced in as its token id, so a pair like AB cannot be split by its neighbours. A noul is a two-option choice whose options read "no" and "yes"; its answer is pyesp_\text{yes}. A score lists its levels as 0: …, 1: … and returns the expected level ∑kk pk\sum_k k\,p_k.

This is the Decider layout, and the card says so. The NOTICE.md credits Decider by Mark Marosi for "the answer-slot prompt layout, the single-token option codes … and the temperature as a function of the number of options". Remember that credit; it matters for the benchmark.

Agr-flash has no chat format to borrow, so it uses a trained head. Its layout wraps the state in a <doc> marker and each option in <choice>…</choice>, and ends each question with <answer>. The hidden state at <answer> becomes a query, each option's tokens are averaged into a key, and the logits are scaled dot products (measured, Scorer.forward()):

pk=softmax⁡k ⁣((WKhˉk)⊤(WQhanswer)256)p_k = \operatorname{softmax}_k\!\left(\frac{(W_K \bar h_k)^\top (W_Q h_\text{answer})}{\sqrt{256}}\right)

WQW_Q and WKW_K are 256 by 960. Two of them plus five 960-wide markers is the 496,320-parameter head (reasoned: 2×256×960+5×9602 \times 256 \times 960 + 5 \times 960). It is the same pointer idea as Kev's, except that here each option is pooled from its own span rather than read at one token, and Agr-flash has no fitted temperature in its config at all.

One pass for every question

The llama.cpp endpoint answers a three-question request with three prompts; with a free slot it copies the state's cache, and without one it reads the state again. Agr puts the whole request in one sequence and changes the attention mask so that it behaves like separate prompts (measured, Layout.visibility()):

So the state is read once, every question is a branch hanging off it, and the answer is read at the last token of each branch. Gemma's sliding-window layers (50 of the 60, window 1,024, measured from layer_types) get the same mask, further limited by position distance; the code builds a separate mask for each layer type and relies on SDPA to apply it. Toggle the plain causal layout to see what the mask removes, and change the sizes to see what reading the state once saves:

who can see whom: one request, one sequence, a branch per question

row = token reading · column = token read · top numbers = position ids

Every question block reads the state and itself. Its position ids restart at 6, as if it were the only question, so adding, removing or reordering questions leaves each answer alone.

tokens through the model (arithmetic)one pass, 3 branches: 2,180one prompt per question, nothing shared: 6,180
The mask is the one agr/model.py builds: the state is causal, each question block sees the state and its own earlier tokens, and its positions start where the state ends. Agr reads the answer at the last token of each block. The token counter is arithmetic, not a timing; on Gemma's sliding-window layers the same mask is further limited to keys within 1,024 positions.

The cost difference is plain arithmetic. With a state of SS tokens and kk questions of qq tokens each, one pass reads S+kqS + kq tokens; kk separate prompts with nothing shared read k(S+q)k(S + q). For a 2,000-token state and three 60-token questions that is 2,180 against 6,180 (reasoned). The design note's claim is that the answers do not change either way: in fp32 an answer is the same whichever other questions come with it "to within 0.00001", and bf16 on a GPU moves answers "by up to about 0.02" (reported; I did not run it). The design is credited to TypeSafe's Jev as described in Archer Hume's "Jev's Architecture Unmasked", and to Kev.

What the mask does not remove is option order. Within one question the options are still read in sequence, and A and B still carry different priors. The code can read each question twice with the options reversed and average the two (order_averaging, credited to Invergent AI's Surogate), but Agr's config.json does not set it (measured), and the card's limitations say it outright: "Changing the order of options can change the answer."

A server caps a request at 64 questions and 1 MB of body; a state or question over its token budget is cut, and the response says "truncated": true; a request over the total is refused with a 422 (measured, agr/serve.py).

Where the probabilities come from, and how they were tuned

"A probability for each option" is a claim about calibration: that 0.8 means right about 80% of the time. Agr's only calibration mechanism is the temperature. Its config.json stores

"temperature_by_options": { "a": 1.25, "b": -0.05, "min": 0.05 }

and the code computes T(n)=max⁡(0.05, 1.25−0.05ln⁡n)T(n) = \max(0.05,\ 1.25 - 0.05 \ln n), "fitted by log loss on our development data" (reported, DESIGN.md). That gives 1.215 for a yes/no question, 1.181 for four options and 0.973 for 255 (reasoned). A temperature above 1 flattens the distribution; Agr's raw logits are slightly overconfident for small nn.

The comparison that makes this interesting is Decider's own checkpoint of the same model. Mapika/decider-chat-gemma4-31b is "stock Gemma-4-31B-it" read through the same layout with "unmodified weights" and its own fitted curve (measured, its decider_config.json):

"temperature_by_options": { "a": 10.124, "b": -1.633, "min": 0.05 }

That is 8.992 at two options, 7.86 at four and 4.804 at twenty-six (reasoned). Stock Gemma's letter logits are so sharp that a yes/no answer has to be divided by nearly nine before its probabilities mean anything. After Agr's fine-tune, the same output layer needs about 1.2. The fine-tune did not just change which answer wins; it trained the hidden state toward logits that are already near the right scale, which is what training on log loss over the option codes does (reasoned; the training recipe is not published).

the fitted temperature, T(n) = max(min, a + b ln n), for two checkpoints of the same Gemma
024681024102664255

x: options (log) · y: temperature · Agr · decider-chat, stock weights

top option ahead of 3 equal rivals by the gap (illustrative)

T = 1 (raw softmax)
top p 0.948 · confidence 0.931
Agr, T(4) = 1.181
top p 0.908 · confidence 0.877
stock Gemma readout, T(4) = 7.860
top p 0.357 · confidence 0.142
A temperature divides the option logits before the softmax. Agr's fitted curve sits between 1.215 at two options and 0.973 at 255; the stock-weights readout of the same Gemma needs 8.992 at two options. Confidence is System One's (p_max − 1/K) / (1 − 1/K), not an accuracy.

Two cautions. First, a single curve per option count corrects average overconfidence; it cannot fix a model that is overconfident on one kind of question and underconfident on another. Second, Agr publishes no calibration number. The Decision Index measures one for every board entrant (a 1-in-6 sample of 32 benchmarks, confidence as the probability on the chosen option): Jev's expected calibration error is 0.074, decider-chat on stock Gemma 0.047, and pplx-decider-v1-27b 0.0178 (measured, from data/index.json). Agr is not on the board, so it has no such figure. The card's own advice is the right one: check thresholds on your own labelled data, which agr eval reports as accuracy, Brier score and calibration error.

The confidence field in a response is not a calibrated probability either. For a choice it is System One's (p_\max - 1/K) / (1 - 1/K): 0 for a uniform spread, 1 for certainty. On four options a top probability of 0.7 is a confidence of 0.6 (reasoned).

Checking the 58.15

Bar chart titled Decision Index, Decision Index 0.2.1: Agr 31B 58.1, Jev (size not disclosed) 57.9, Rune 26B-A4B v3 at 25.8B 57.4, Kev 9B at 9.7B 38.5. Footer: Agr: our own run of the public kit. Others: the public board, 28 September 2026.
The launch chart. Agr's bar is Command Code's own run; the other three are copied from the public board of 28 September. The board's second place, the same Gemma checkpoint with no training, is not shown. (Command Code, Agr model card.)

The headline holds as arithmetic. The Decision Index 0.2.1 is a weighted mean of five areas' chance-corrected skill: Knowledge 25.8%, Language 25.8%, Retrieval 20.0%, Tools 18.3%, Arts 10% (measured, index02.py in the kit at 87d4650). Applying those weights to the board's area scores for Jev gives 57.91, its board figure. Applying them to Agr's five rounded area scores from its own chart gives 58.14, consistent with 58.15 after rounding (reasoned). One reply asked whether 58.1 on the chart and 58.15 in the post disagree; they do not, the chart rounds to one place. The "150,759 requests" on the card is the kit's 120,340 rescored rows plus the 30,419 added in 0.2 (measured, the kit's README and tests).

What the number is not is a board entry. The footer says "Agr: our own run of the public kit". The board's index.json was generated on 28 September and has 70 entrants and no Agr (measured). A reply asked whether Agr will be scored on the public board; nothing says so yet. A quarter of a point over Jev (58.15 against 57.91, reasoned) from a self-run is a tie until someone else runs it.

Table titled Results by area. Columns Agr 31B, Jev, Rune v3 25.8B, Kev 9.7B. Decision Index 58.1, 57.9, 57.4, 38.5. Knowledge 45.8, 51.4, 43.4, 26.2. Language 59.5, 62.0, 63.1, 41.7. Retrieval and routing 65.3, 55.4, 63.5, 43.7. Tools 75.9, 75.1, 71.2, 54.5. Arts and judgement 39.7, 37.7, 41.9, 22.4.
Agr's areas against three board entrants. Agr wins Retrieval by ten points over Jev and loses Knowledge by 5.6. (Command Code, Agr model card, results by area.)

Per area, Agr is ahead of Jev on Retrieval (65.3 against 55.4), Tools (75.9 against 75.1) and Arts (39.7 against 37.7), and behind on Knowledge (45.8 against 51.4) and Language (59.5 against 62.0) (reported for Agr, measured for Jev). The Knowledge gap is reasoning: GPQA Diamond 34.0 against 71.4, BBH 66.5 against 89.7 (same labels). Across the 38 benchmarks Agr is ahead of Jev on 21 and behind on 17 (reasoned, from the per-benchmark chart below). The Jev values in Command Code's charts match the board file to the decimal I checked (measured).

The comparison the chart leaves out

The board's second place, at 57.33, is "Decider chat · Gemma-4-31B": stock google/gemma-4-31B-it, no training, read through Decider's chat layout with the 10.124 temperature curve (measured, data/index.json and the decider-chat config). That is Agr's backbone and Agr's readout without Agr's fine-tune. It is the control experiment, it was on the board Command Code copied its other bars from, and the chart shows Jev, Rune and Kev instead.

Put Agr's area scores next to it (reported for Agr, measured for decider-chat, differences reasoned):

AreaWeightAgrSame Gemma, no trainingChange
Knowledge25.8%45.844.3+1.5
Language25.8%59.560.4−0.9
Retrieval20.0%65.363.1+2.2
Tools18.3%75.975.6+0.3
Arts10%39.738.3+1.4
Index58.1557.33+0.82

The fine-tune is worth 0.82 index points over the same weights read the same way. That is real; 26 of 38 benchmarks go up, 10 go down and 2 do not move (reasoned). But it is not where the launch post says the effort went. The Tools area, five benchmarks, moves 0.3: API-Bank 84.8 to 85.0, BFCL 97.3 to 97.3, Home appliances 62.5 to 63.6, ToolRet 58.0 to 60.3, When2Call 69.2 to 67.2. The big gains are elsewhere: Amazon ESCI and HoVer up 7.3 each, SATA-Bench 7.2, ForecastBench 9.1. The big losses are NLI4CT down 9.5, PhishNChips down 7.5 and iSarcasmEval down 7.1.

Table titled Results by benchmark, 38 rows grouped into Knowledge, Language, Retrieval and routing, Tools, and Arts and judgement, with columns Agr, Jev, Rune and Kev and the best value in each row highlighted. Tools rows: API-Bank 85.0, 88.0, 83.0, 55.5; BFCL 97.3, 94.3, 93.0, 92.6; Home appliances 63.6, 52.3, 46.6, 25.0; ToolRet 60.3, 59.9, 58.9, 58.7; When2Call 67.2, 74.6, 68.0, 32.8.
All 38 counted benchmarks. Agr's highlighted wins are real against these three columns; against the untrained Gemma readout most of the Tools rows are within two points. ForecastBench is labelled 'Brier (lower is better)' but shows skill, where higher is better, as the card's text says. (Command Code, Agr model card, results by benchmark.)

"Best-in-class on coding agent tool calls and routing" needs two corrections (reasoned, from the board). The Decision Index has no coding-agent benchmark, so that part cannot be checked against it at all; the Tools area is general function calling and API use. And on that area, three board entrants already score higher than Agr's 75.9: pplx-decider-v1-27b at 79.3, Jebadiah 27B at 78.1 and simple-jev on Qwen3.8-27B at 76.2 (measured). Where Agr does lead every board entrant is the area its chart calls "Retrieval and routing": 65.3 against Rune's 63.5. If "routing" means that, the claim holds there.

Agr-flash, unscored

Several replies asked for Agr-flash's number. There is none. The card gives it a row in the models table, token limits and the line "Agr-flash is not meant for safety decisions", and every chart is Agr. A 362M model with a trained head is the interesting deployment shape, the one the replies kept pointing at, and nothing published says how much it gives up. For scale, the 0.8B Kev on the same board scores 14.6 (measured); I would not guess where Agr-flash lands.

Smaller things the files say

What holds

The mechanism is clean and documented: one pass over the state, a masked branch per question, Gemma's own output layer restricted to the option codes, a fitted temperature, and a design note that says which pieces came from Decider, Jev, Kev and Surogate. Both checkpoints are what the card says, at the sizes it says, under Apache 2.0, and the fine-tune is visible in the bytes. The temperature comparison is the clearest evidence of what training a decision model does to a chat model: the same output layer goes from needing a divisor of about nine to needing about 1.2.

The 58.15 is a self-run quarter-point over Jev, and the chart omits the one bar that shows how much of it is the fine-tune: 0.82 points over the same Gemma read the same way, with almost none of it in Tools. If you are choosing a decision model for tool routing, the Decision Index's Tools column says Agr is in the top group and not at the top of it, and the card's last advice applies to every model in this category: measure on your own labelled requests before trusting a threshold.

Sources, read on 6 October 2026: Command Code's post, its thread and first page of replies through the fxtwitter mirror; CommandCode/agr and CommandCode/agr-flash (cards, configs, NOTICE.md, safetensors headers and byte ranges); CommandCodeAI/agr at 296303d (agr/model.py, agr/serve.py, DESIGN.md, pyproject.toml); google/gemma-4-31B-it and HuggingFaceTB/SmolLM2-360M byte ranges; Mapika/decider-chat-gemma4-31b decider_config.json and Mapika/decider; the Decision Index kit at 87d4650 and the board's data/index.json (generated 28 September 2026) from the Decision Index space.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Agr: a decision model is Gemma's own output layer, read at one position per question", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026agrdecisionmodel,
  author = {Satyajit Ghana},
  title  = {Agr: a decision model is Gemma's own output layer, read at one position per question},
  url    = {https://ai.thesatyajit.com/articles/agr-decision-model},
  year   = {2026}
}
share