2026-10-02 · 18 min · explainer · llm · ecosystem · benchmarks · calibration · architecture
Week one was code with almost no numbers. Week two brought seven open-weight decision models with numbers, and the finding was that most of the comparisons stood on JevBench's 231 public items, that on the 308 sealed items a three-point public lead over Jev turned into a 2.9-point deficit, and that two releases out of twelve measured Jev themselves and printed the loss.
Week three is seven more posts, and the releases are, by some distance, the most serious in the category so far. The architectures stopped being merged LoRAs with the vision tower left untouched: this week has a trained 128-million-parameter cross-attention head from Cloudflare, a readout over 62 symbols from Shanghai AI Lab, a serving layer that makes the whole trick a flag, and a from-scratch 1.2-million-parameter model that does not speak text at all.
The evaluation went the other way. Not one of this week's new models carries an independent or sealed-item JevBench score. clef and Lumma-fev publish their own suites and never touch it; Intern-Decision runs its own copy of the public tiers; and JEMM's lone JevBench row is a loss it files under "gap not significant." Week two put two releases on the sealed board. Week three: none.
Parameter counts below are safetensors header or metadata sums, not the cards' round numbers. Code and configs were read, never run. Measured means read from a file or recomputed from published rows; reported means the author's claim, which I could not check; reasoned means my inference. All checked on 2026-10-02.
The one idea, again
A typed decision model does one thing a chat model behind a harness does not: it answers a schema of typed questions — choice, score, yes/no — by returning a calibrated probability for every option in a single forward pass, with no text generated and nothing to parse. A chat model writes a label into JSON and you hope the format holds; a decision model reads the answer off the logits, or off a trained head, at the position where the answer goes.
Week two sorted the open ones into three families by how they turn options into probabilities, and the same split holds this week:
- Letter (vocabulary) readout. Options go into the prompt as letters, and
the answer is the logits of those letter tokens.
AandBcarry different priors, so order matters, and the alphabet caps the option count. - Pointer. Options are spans in one shared sequence, and a trained head points at one. Options see each other; order still matters. This is Kev's family.
- Joint / per-option scorer. Each option gets a learned scalar. In the isolated version no option sees another and order cannot matter; in the joint version a trained head scores all options of all questions together.
From each release's card, config and head tensors (measured): Intern-Decision and JEMM are letter readouts. Lumma-fev is a pointer. clef is the new one — a trained joint scorer. taiga-s1 is a per-command scorer for a domain that has nothing to do with language. SGLang is not a model; it is the serving layer that makes a letter readout out of anything. Laya on Unsloth is a packaging of an existing model.
Jev's median single-request latency, as four publishers measured it (ms, log axis)
- 106.3 ms — Intern-Decision (RTX 4090, local HF)
- 256 ms — Lumma-fev (card P50)
- 455 ms — JEMM (paired, short)
- 524.1 ms — clef (Decision Index median)
Intern-Decision Shanghai AI Lab
new model- size
- 0.85B / 2.2B / 4.5B params
- modality
- text + up to 8 images
- decides by
- symbol readout (A–Z, a–z, 0–9)
- option cap
- 62 options, 16 questions
- vs Jev
- 4B avg 90.02 vs Jev 88.74 across seven benches (self-run)
- calibration
- 4B ECE 0.065 vs Jev 0.095 (self-run)
- license
- Apache-2.0
JEMM MaestroYan
new model- size
- 117M LoRA on Qwen3.8-27B
- modality
- text + screenshot
- decides by
- letter readout (A–Z, 0–5)
- option cap
- 32 options
- vs Jev
- wins its own held-out splits; JevBench public 198 vs 200 of 231 (−2)
- calibration
- none published
- license
- Apache-2.0
Lumma-fev FrontiersMind
new model- size
- 0.15B / 0.6B / 4B / 9B params
- modality
- text
- decides by
- Kev-style pointer head
- option cap
- 255 options
- vs Jev
- 4B avg 0.90 vs Jev 0.75 on four tasks; Jev column copied from Laya's card
- calibration
- none published
- license
- Apache-2.0
clef / clef-flash Cloudflare
new model- size
- 27.4B + 128M head / 9.4B + 122M head
- modality
- text, JSON, images, video
- decides by
- trained joint schema head
- option cap
- schema-scored
- vs Jev
- wins most of its own Decision Index; Jev wins the reasoning benches; no JevBench
- calibration
- ForecastBench Brier 13.9 vs Jev 17.4 (self-run)
- license
- Apache-2.0
SGLang /v1/decisions SGLang
serving layer- size
- any served chat model
- modality
- text (VLM in the demo)
- decides by
- label logprobs via /v1/score
- option cap
- 26 (255 with two-letter labels)
- vs Jev
- /v1/systemone speaks the TypeSafe SDK; no accuracy claim in the docs
- calibration
- label_mass only; docs say it is not calibrated
- license
- Apache-2.0 (SGLang)
Laya on Unsloth Unsloth
packaging- size
- 678–846 MB (the Laya model)
- modality
- text
- decides by
- Laya, unchanged
- option cap
- Laya's
- vs Jev
- runs Laya locally behind a Jev-compatible API; no new numbers
- calibration
- Laya's own
- license
- Laya + Unsloth
taiga-s1 shhivv
oddball- size
- 1.2M params, from scratch
- modality
- FreeCAD state (no text model)
- decides by
- per-command scorer
- option cap
- available commands
- vs Jev
- not a Jev competitor: a ~1 ms FreeCAD action decider
- calibration
- held-out ECE 0.0272 (measured, in config)
- license
- MIT
The new models
clef (Cloudflare): the first one that trained a real head
The post is a one-liner — "another open source Jev, this time by Cloudflare,
post-trained on Qwen 3.8" — and it undersells the release. Cloudflare/clef is
the most complete thing in the category so far.
- architecture
- Qwen3_5ForConditionalGeneration
- task
- image-text-to-text
- library
- transformers
- license
- apache-2.0
- safetensors
- 13 shards
- largest file
- 4.99 GB
- files
- 25
- downloads
- 18
- likes
- 458
A full 27,356,728,560-parameter Qwen3.8-27B backbone (BF16, twelve shards, vision encoder included) plus a separate joint_head.safetensors of 128,056,324 parameters, counted from the safetensors headers by range request. clef-flash is a 9,409,813,744-parameter backbone with a 121,762,820-parameter head. Apache-2.0, following the base model.
repo last modified 2026-10-01
What makes it different is the head, and I read it tensor by tensor (measured,
joint_head.safetensors, joint_head_config.json). It is not a letter readout
and not a LoRA. It is a 128-million-parameter cross-attention transformer:
two "evidence" layers that route information from the state to each question,
then four transformer layers, projections pulling the backbone's 5,120-wide
hidden states down to a 1,024-wide working space, and a residual scorer MLP. The
card's own words: it "scores all options of all questions jointly." So unlike
every isolated scorer before it, clef's options do see each other — the first
head in this category trained as a unit rather than folded in as an adapter. It
also reads images and video, where Jev-Omni first pushed
the category multimodal.
The benchmark table is the biggest of the week — about 42 rows — and it is also the catch. Every number is from "our internal run of the Decision Index 0.2.1 suite" (reported), Cloudflare's own benchmark. clef wins most of it: BANKING77 macro-F1 94.2 against Jev's 79.7, CLINC150 97.4 against 89.3, ForecastBench Brier 13.9 against Jev's 17.4 (lower is better). But read the rows where Jev wins, because they are the reasoning-heavy ones: GPQA Diamond 78.3 against clef's 48.0, MMLU-Pro 82.7 against 65.9, BBH 92.9 against 73.7. The pattern (reasoned) is clean: clef wins routing, classification and structured extraction; Jev wins the benches that need actual inference. Median latency on the same suite: clef 209.3 ms, clef-flash 38.8 ms, Jev 524.1 ms (reported).
None of this is JevBench. clef is absent from the one third-party board the rest of the category reports on, so there is no way to rank it against the week-two releases: the best-documented card of the wave is the one you can place on no independent axis.
Intern-Decision (Shanghai AI Lab): 62 symbols, and a fitted temperature
The post, from ModelScope, is specific: three sizes averaging 79.38, 84.68 and 90.02 across seven decision benchmarks, the 4B "surpasses Jev at 88.74 while achieving better probability calibration," and mean latency **33.98, 33.28 and 44.16 ms against 109.70 ms for Jev in the same local HF setup."
- architecture
- Qwen3_5ForConditionalGeneration
- task
- image-text-to-text
- library
- transformers
- license
- apache-2.0
- safetensors
- 4 shards
- largest file
- 4.26 GB
- files
- 21
- downloads
- 1.3K
- likes
- 73
Base Qwen3.5-4B. Safetensors totals (measured): 0.8B is 852,985,920, the 2B 2,213,241,664 and the 4B 4,539,265,536 parameters. Apache-2.0, with the Qwen3.5 upstream notices retained. A separately fitted candidate-probability temperature of 1.99241824 ships with the 4B.
repo last modified 2026-09-26
The mechanism is a letter readout with two refinements (measured, from the card's
inference section). It maps each question's options to single-token symbols —
A–Z, then a–z, then 0–9, 62 of them — so the cap is 62 options
where XOR and JEMM stop at 26. It renders a full JSON answer skeleton with one
<decision> placeholder per field, runs one forward pass, and reads the logits
at the position immediately before each placeholder, so up to 16 questions are
answered in that single pass. Then it applies a fitted candidate-probability
temperature — softmax(log(p) / T) with T = 1.99241824, tuned by NLL
minimisation on 1,728 held-out cases — which "updates confidence while preserving
the argmax decision." It is still a vocabulary readout, so A and B carry
different priors and order still matters; the card does not mention reversing and
averaging.
The arithmetic checks (measured): the 4B's seven scores are 100.00, 98.61, 73.87, 80.55, 96.45, 90.82 and 89.86, which average to 90.02, and Jev's seven average to 88.74. So "surpasses Jev at 88.74" means the 4B's 90.02 beats Jev's 88.74 — the number in the headline is Jev's. On calibration the 4B reports ECE 0.065 against Jev's 0.095 and Brier 0.347 against 0.358, and a separate 96-case pilot puts the 4B at ECE 0.089 after calibration against Jev's 0.130 (all reported).
Two catches. Three of the seven benchmarks are JevBench-Easy, -Original and -Hard — the public tiers, the 48 + 72 + 111 that make up the 231 public items — run by internlm's own harness, not JevBench's. And the Jev comparison, accuracy and latency both, is internlm running Jev itself: the latency table says "Jev in the same local HF setup" at 109.70 ms, but Jev is TypeSafe's hosted, closed model, and the card never explains how it ran locally on a 4090. I could not verify that number, or that the Jev benchmark rows are the hosted model's.
JEMM (MaestroYan): the candid figure, and the loud post
The post, in Chinese, leads with "85.7% crushes Laya, tool-calling 93.3%
overtakes Jev." MaestroYan/JEMM is a LoRA adapter (measured,
adapter_config.json): r=16, α=32, 116,727,808 parameters, no head tensor at
all, on Qwen/Qwen3.8-27B. It is a pure letter readout — labels A–Z and
0–5, 2 to 32 candidates — that reads the base model's own last-token logits
over those label tokens and softmaxes them. Multimodal, because Qwen3.8 is: it
takes a screenshot (trained at 1280×800). Its training data is listed —
Mind2Web, BFCL, ToolACE, Banking77, CLINC150 and more — and "no Jev outputs were
used."

The figure is admirably honest and the post is not. The six wins JEMM charts big are on its own held-out sets: web actions on Mind2Web (36.0 to 54.6 with a screenshot, Jev given only text; n=894), tool calling on BFCL (88.7 to 93.3; n=2,148), tool abstention on MetaTool (92.6 to 99.3). Those are real, paired, with lower bounds. But the one row on an outside benchmark, JevBench public, is 198 of 231 against Jev's 200 (measured, from the figure) — a loss, filed under "on par, or gap not significant." That is 85.7%, the post's headline number — above Laya's 58.4%, below Jev's. There is no sealed-tier number, and no calibration number anywhere.
The latency figure is as careful and has the same shape: single-question medians of 271, 274 and 299 ms against Jev's 455, 454 and 458 — 0.60× to 0.65×, a −31.2% aggregate across the five text configs (measured, from 600 paired requests). But JEMM runs on your GPU and Jev is a hosted API, so "0.60× of Jev's time" is partly the network round-trip to TypeSafe, not compute — local-versus-hosted is the most common unstated confound in this category's latency claims.
Lumma-fev (FrontiersMind): the training write-up, the same copied column
Week two covered Lumma-fev's 0.1B and 0.6B and noted the 4B and 9B were promised but not on Hugging Face. They are now, and the post this week is the training write-up.
Measured, from the safetensors metadata and config.json: the 4B is
4,207,062,528 parameters with a 1,311,232-parameter FP32 head — exactly
Kev-4B's pointer-head size, the same count NeoHorse-Jev-4B
borrowed — and the 9B is 7,938,782,208 with a 2,097,664-parameter head. The
config sets option_isolation: false, so it is a pointer: options attend to each
other inside a question, and order matters. The blog adds the training recipe,
which is the genuinely new content: the 4B is continual pre-training on the
Qwen3.5-4B line, then decision fine-tuning; the 9B the same on the 9B line; the
0.6B is a frozen backbone with a LoRA; and the 0.1B is a from-scratch
Nandi-Mini-150M backbone, fully fine-tuned.
What the blog does not add is a re-measured Jev. The card's four-benchmark table (measured, arithmetic recomputed) puts the 4B at an average of 0.90 and Jev at 0.75 across Typed-decisions, AG News, DAIR Emotion and Banking77, with GLiNER-2.5-Decide fourth at 0.54 — but the Jev column (0.72, 0.91, 0.48, 0.87) is identical, to the digit, to the figures on Laya's own model card that week two already flagged, including DAIR Emotion's implausibly low 0.48 for a frontier decision model. Only one of the four, Typed-decisions, is decision-native, and there the 4B leads 0.78 to 0.72. No n, no eval code, no calibration. See Laya vs Jev for the background on the copied column.
The serving layer: SGLang makes the readout first-class
The most consequential post of the week is not a model. SGLang added
/v1/decisions and /v1/systemone, and that is the any model can be
Jev thesis shipped as product.
That piece found the System One readout was a serving feature fifteen months older
than the category: one prefill, max_new_tokens: 0, read the logprobs of the
tokens you name. /v1/decisions (measured, from the docs) is that, wrapped: you
send an input and typed questions, and "each answer comes back with the
probability of every option, read from the model's next-token scores at the answer
position. No text is generated and no output is parsed. It needs no special
checkpoint." Choice questions take 2 to 26 options labelled A–Z, score
questions 2 to 10 levels labelled 0–9, one prefill each — a letter readout,
with all its limits.
Two details are better than most of the models'. The response carries
label_mass, "the full-vocabulary probability of the answer labels at the
answer position. A low value means the model puts most of its probability outside
the offered answers" — exactly the abstention signal the any-model piece said
renormalisation throws away, surfaced as a first-class field. And the docs are
blunt: "None of these values is a calibrated probability that the decision is
correct. Validate any threshold on labeled data from your workload." What they do
not touch is the batching non-determinism Jev is not
deterministic found — a scored answer read
off the logits can still move between batch sizes, serving layer or not.
/v1/systemone is the compatibility layer: it "serves the same decisions in the
request and response shape of the System One API," so "the official TypeSafe
SDKs" work by pointing their base URL at the server. It reaches Jev's full 255
options, but only by handing every option past 26 a two-letter label (AA, AB,
…) that the docs warn "include common words and have unequal priors, so answers
above 26 options can depend on option order." The demo attached to the post —
Qwen3.8-27B turned into a multimodal player that cleared Pokémon FireRed's Elite
Four with sub-100 ms decisions — is a reported showcase, not in the docs. This is
the infrastructure that makes the models above interchangeable: one API, any
backbone, no fine-tune — and no calibration guarantee either.
Packaging: Laya on Unsloth
Unsloth's post is a deployment story, not a model: "run Laya Decision models locally on just 4GB RAM," on CPU, Mac, Windows, Linux or GPU, "through a Jev-compatible API via Unsloth Desktop." The docs (reported) give the real shape: the default multilingual Laya is a 678 MB model needing 4 GB of RAM with a 1,024-token context, the English and typed-decision variants are 846 MB; CPU is the default, the GPU is optional and falls back to CPU on an out-of-memory. The Unsloth GUI "expose[s] Laya through a TypeSafe-compatible Jev API, so existing Jev integrations can work by pointing them at your local server."
There is nothing new to benchmark here — Laya is the same model Laya vs Jev and Laya on MLX covered — and that is the point. The "4 GB RAM" is the host requirement, not the model; a 678 MB encoder classifier was already small. What Unsloth adds is the one-click local server speaking Jev's wire format, which is the same thing localjev, lumma-fev-serve and now SGLang each built independently. The wire format is by now reproduced about as many times as the readout.
The oddball: taiga-s1 builds parts in FreeCAD
shhivv/taiga-s1 is the one that does not fit the frame, and it is the most fun.
A 1,228,163-parameter model (measured, safetensors total — "1.2M"), trained
from scratch, no language model and no vision model, that builds 3D parts in
FreeCAD. "Your agent plans, taiga executes": a planner hands it an ordered
feature list ("plate 40×30×10, Ø6 hole, polar pattern ×6, fillet the top edges")
and taiga picks the next FreeCAD command, step by step, from the ones currently
available, at about 1 ms per decision on a CPU (reported).
It is a decision model in the strict sense this series uses: a 3-layer encoder
and 2-layer decoder (measured, config.json: S1Model, width 128, 4 heads)
where "candidate commands attend to the state and the active goal item, and each
gets one score" — a per-option scorer over the available actions, with a fitted
temperature and calibrated probabilities. And the calibration is the only one
this week I could read straight from the repo: config.json records a held-out
set of 5,573 states at ECE 0.0272 after temperature scaling (0.0416 at T=1),
accuracy 0.9551, fitted temperature 2.554 (measured).
It makes no claim against Jev and belongs on no JevBench tier. It is here because it is the clearest answer to "what is a decision model, minimally?" — a tiny head that scores a constrained set of typed options and knows how confident it is. You do not need 27 billion parameters for that. You need 1.2 million, if the option set is a CAD workbench rather than the open world.
The evaluation went backwards
Line up what week three measured against what week two did, and the trend is not the one the posts imply.
Nobody is on the sealed board. Week two put XOR and Open-Jev on JevBench's 308 sealed items, where both sat below Jev. Week three's four new models are on no sealed tier: clef and Lumma publish their own suites, Intern-Decision runs its own copy of the public tiers, and JEMM's only JevBench number is a public one it loses. The sealed items were the whole point of week two's finding — a public lead of a few points said nothing about the sealed ranking, and reproducing Jev showed how quickly a board fills with rows aimed at the public items — and this week not one new release went near them.
The baseline is not a constant. Jev is a single closed model, and this week's publishers report its median single-request latency as 106.30 ms (Intern-Decision, on a 4090 "local HF"), 256 ms (Lumma's card), 455 ms (JEMM's paired run) and 524.1 ms (clef's Decision Index). That is a 5× spread on the same model, because each ran it a different way — some against the hosted API with its network hop, some in a local configuration I cannot reconstruct for a closed model. A "−31.2% vs Jev" is a statement about the publisher's stopwatch as much as about the model.
Calibration is still the tell. The one property you cannot fake from the serving layer is the one that separates the careful releases. taiga-s1 ships a measured ECE in its config. Intern-Decision reports ECE and Brier and a calibration pilot. clef reports a ForecastBench Brier. JEMM and Lumma-fev publish none, and SGLang and Unsloth say plainly that what they return is not calibrated and you must measure your own. That split — who publishes a reliability number — is still, as it was in any model can be Jev, the fastest way to tell who did the expensive half.
If you are choosing this week
All reasoned, from the evidence above:
- A real trained head, documented: clef — the first joint scorer trained as a unit, with the fullest card. Every number on it is Cloudflare's own run, on no third-party board.
- A calibration temperature that ships: Intern-Decision and taiga-s1 both fit one and tell you the value; taiga's held-out ECE is in the repo.
- More than 26 options: Lumma-fev's pointer (255) or Intern-Decision's 62 symbols. On SGLang you reach 255 only through order-dependent two-letter labels.
- Drop-in on your own backbone: SGLang
/v1/systemone, no fine-tune — and no calibration guarantee. - A sealed-item or independent JevBench score: nothing this week. The week-two releases on the board are still the only ones with one.
The take
Week three is the strongest set of open decision models yet — a trained joint head, a 62-symbol readout, a serving layer that makes the whole thing a flag, a 1.2-million-parameter CAD decider — and the weakest set of evaluations. Every "beats Jev" number is the publisher's own, most are not on JevBench at all, none is sealed, and the Jev they are beating has a latency that depends on who held the stopwatch.
The one release that measured a reliability number straight into its config file is the 1.2-million-parameter one that does not claim to beat anything. That is the week, in one sentence.
Sources, read on 2026-10-02: Cloudflare/clef and Cloudflare/clef-flash; internlm/Intern-Decision-4B, -2B and -0.8B; MaestroYan/JEMM; the Lumma-fev collection and its training write-up; shhivv/taiga-s1 and its code; SGLang's decision-models docs; and Unsloth's Laya + Jev API guide. Parameter counts and head architectures are safetensors header and metadata reads; no third-party code was executed. The figure is reproduced for commentary from JEMM's model card (Apache-2.0); the interactive is original.