# Jev alternatives, week three: bigger models, smaller benchmarks

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/jev-alternatives-week-three
> date: 2026-10-02
> tags: explainer, llm, ecosystem, benchmarks, calibration, architecture

[Week one](/articles/jev-ecosystem) was code with almost no numbers.
[Week two](/articles/jev-alternatives-week-two) brought seven open-weight
decision models with numbers, and the finding was that most of the comparisons
stood on JevBench's **231 public items**, that on the **308 sealed items** a
three-point public lead over Jev turned into a 2.9-point deficit, and that two
releases out of twelve measured Jev themselves and printed the loss.

Week three is seven more posts, and the releases are, by some distance, the most
serious in the category so far. The architectures stopped being merged LoRAs with
the vision tower left untouched: this week has a trained 128-million-parameter
cross-attention head from Cloudflare, a readout over 62 symbols from Shanghai AI
Lab, a serving layer that makes the whole trick a flag, and a from-scratch
1.2-million-parameter model that does not speak text at all.

The evaluation went the other way. **Not one of this week's new models carries an
independent or sealed-item JevBench score.** clef and Lumma-fev publish their own
suites and never touch it; Intern-Decision runs its own copy of the public tiers;
and JEMM's lone JevBench row is a loss it files under "gap not significant." Week
two put two releases on the sealed board. Week three: none.

Parameter counts below are safetensors header or metadata sums, not the cards'
round numbers. Code and configs were read, never run. **Measured** means read
from a file or recomputed from published rows; **reported** means the author's
claim, which I could not check; **reasoned** means my inference. All checked on
2026-10-02.

## The one idea, again

A [typed decision model](/articles/any-model-can-be-jev) does one thing a chat
model behind a harness does not: it answers a schema of typed questions —
choice, score, yes/no — by returning a calibrated probability for every option
in a **single forward pass**, with no text generated and nothing to parse. A
chat model writes a label into JSON and you hope the format holds; a decision
model reads the answer off the logits, or off a trained head, at the position
where the answer goes.

[Week two](/articles/jev-alternatives-week-two) sorted the open ones into three
families by *how* they turn options into probabilities, and the same split holds
this week:

- **Letter (vocabulary) readout.** Options go into the prompt as letters, and
  the answer is the logits of those letter tokens. `A` and `B` carry different
  priors, so order matters, and the alphabet caps the option count.
- **Pointer.** Options are spans in one shared sequence, and a trained head
  points at one. Options see each other; order still matters. This is
  [Kev](/articles/kev)'s family.
- **Joint / per-option scorer.** Each option gets a learned scalar. In the
  isolated version no option sees another and order cannot matter; in the joint
  version a trained head scores all options of all questions together.

From each release's card, config and head tensors (measured): **Intern-Decision**
and **JEMM** are letter readouts. **Lumma-fev** is a pointer. **clef** is the
new one — a trained joint scorer. **taiga-s1** is a per-command scorer for a
domain that has nothing to do with language. **SGLang** is not a model; it is the
serving layer that makes a letter readout out of anything. **Laya on Unsloth** is
a packaging of an existing model.

<EntrantMatrix />

## The new models

### clef (Cloudflare): the first one that trained a real head

The post is a one-liner — "another open source Jev, this time by Cloudflare,
post-trained on Qwen 3.8" — and it undersells the release. `Cloudflare/clef` is
the most complete thing in the category so far.

<ModelCard repo="Cloudflare/clef" note="A full 27,356,728,560-parameter Qwen3.8-27B backbone (BF16, twelve shards, vision encoder included) plus a separate joint_head.safetensors of 128,056,324 parameters, counted from the safetensors headers by range request. clef-flash is a 9,409,813,744-parameter backbone with a 121,762,820-parameter head. Apache-2.0, following the base model." />

What makes it different is the head, and I read it tensor by tensor (measured,
`joint_head.safetensors`, `joint_head_config.json`). It is not a letter readout
and not a LoRA. It is a **128-million-parameter cross-attention transformer**:
two "evidence" layers that route information from the state to each question,
then four transformer layers, projections pulling the backbone's 5,120-wide
hidden states down to a 1,024-wide working space, and a residual scorer MLP. The
card's own words: it "scores all options of all questions jointly." So unlike
every isolated scorer before it, clef's options *do* see each other — the first
head in this category trained as a unit rather than folded in as an adapter. It
also reads images and video, where [Jev-Omni](/articles/jev-omni) first pushed
the category multimodal.

The benchmark table is the biggest of the week — about 42 rows — and it is also
the catch. Every number is from "our internal run of the Decision Index 0.2.1
suite" (reported), Cloudflare's own benchmark. clef wins most of it: BANKING77
macro-F1 94.2 against Jev's 79.7, CLINC150 97.4 against 89.3, ForecastBench Brier
**13.9 against Jev's 17.4** (lower is better). But read the rows where Jev wins,
because they are the reasoning-heavy ones: GPQA Diamond 78.3 against clef's 48.0,
MMLU-Pro 82.7 against 65.9, BBH 92.9 against 73.7. The pattern (reasoned) is clean:
clef wins routing, classification and structured extraction; Jev wins the benches
that need actual inference. Median latency on the same suite: clef 209.3 ms,
clef-flash 38.8 ms, Jev 524.1 ms (reported).

None of this is JevBench. clef is absent from the one third-party board the rest
of the category reports on, so there is no way to rank it against the week-two
releases: the best-documented card of the wave is the one you can place on no
independent axis.

### Intern-Decision (Shanghai AI Lab): 62 symbols, and a fitted temperature

The post, from ModelScope, is specific: three sizes averaging **79.38, 84.68 and
90.02** across seven decision benchmarks, the 4B "surpasses Jev at 88.74 while
achieving better probability calibration," and mean latency **33.98, 33.28 and
44.16 ms against 109.70 ms for Jev in the same local HF setup."

<ModelCard repo="internlm/Intern-Decision-4B" note="Base Qwen3.5-4B. Safetensors totals (measured): 0.8B is 852,985,920, the 2B 2,213,241,664 and the 4B 4,539,265,536 parameters. Apache-2.0, with the Qwen3.5 upstream notices retained. A separately fitted candidate-probability temperature of 1.99241824 ships with the 4B." />

The mechanism is a letter readout with two refinements (measured, from the card's
inference section). It maps each question's options to single-token symbols —
`A`–`Z`, then `a`–`z`, then `0`–`9`, **62 of them** — so the cap is 62 options
where XOR and JEMM stop at 26. It renders a full JSON answer skeleton with one
`<decision>` placeholder per field, runs one forward pass, and reads the logits
*at the position immediately before each placeholder*, so up to 16 questions are
answered in that single pass. Then it applies a fitted candidate-probability
temperature — `softmax(log(p) / T)` with `T = 1.99241824`, tuned by NLL
minimisation on 1,728 held-out cases — which "updates confidence while preserving
the argmax decision." It is still a vocabulary readout, so `A` and `B` carry
different priors and order still matters; the card does not mention reversing and
averaging.

The arithmetic checks (measured): the 4B's seven scores are 100.00, 98.61, 73.87,
80.55, 96.45, 90.82 and 89.86, which average to 90.02, and Jev's seven average to
88.74. So "surpasses Jev at 88.74" means the 4B's 90.02 beats Jev's 88.74 — the
number in the headline is Jev's. On calibration the 4B reports ECE 0.065 against
Jev's 0.095 and Brier 0.347 against 0.358, and a separate 96-case pilot puts the
4B at ECE 0.089 after calibration against Jev's 0.130 (all reported).

Two catches. Three of the seven benchmarks are JevBench-Easy, -Original and
-Hard — the **public tiers**, the 48 + 72 + 111 that make up the 231 public items
— run by internlm's own harness, not JevBench's. And the Jev comparison, accuracy
and latency both, is internlm running Jev itself: the latency table says "Jev in
the same local HF setup" at 109.70 ms, but Jev is TypeSafe's hosted, closed
model, and the card never explains how it ran locally on a 4090. I could not
verify that number, or that the Jev benchmark rows are the hosted model's.

### JEMM (MaestroYan): the candid figure, and the loud post

The post, in Chinese, leads with "85.7% crushes Laya, tool-calling 93.3%
overtakes Jev." `MaestroYan/JEMM` is a LoRA adapter (measured,
`adapter_config.json`): r=16, α=32, **116,727,808 parameters**, no head tensor at
all, on `Qwen/Qwen3.8-27B`. It is a pure letter readout — labels `A`–`Z` and
`0`–`5`, 2 to 32 candidates — that reads the base model's own last-token logits
over those label tokens and softmaxes them. Multimodal, because Qwen3.8 is: it
takes a screenshot (trained at 1280×800). Its training data is listed —
Mind2Web, BFCL, ToolACE, Banking77, CLINC150 and more — and "no Jev outputs were
used."

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-three/fig1.png"
  alt="A lollipop chart titled 'JEMM vs Jev 1.13 — Accuracy, same sealed questions, paired one by one.' Six rows show JEMM ahead of Jev 1.13: web actions with screenshot 36.0 to 54.6 (+18.6 pp), web actions text only 35.2 to 45.8 (+10.7), content moderation category 54.4 to 63.2 (+8.8), tool abstention 92.6 to 99.3 (+6.7), tool calling 88.7 to 93.3 (+4.6), content moderation 90.2 to 93.0 (+2.8). A lower 'Also measured — on par, or gap not significant' panel has six tiles: batched 4 questions 75 vs 65, batched 8 questions 51 vs 40, Banking77 intent 89.0 vs 87.0, CLINC150 intent 89.5 vs 88.8, MetaTool similar tools 71.2 vs 72.7 (−1.5), and JevBench public 198 vs 200 of 231 (−2)."
  caption="JEMM's own accuracy figure. The six headline wins are on JEMM's own held-out splits (Mind2Web, BFCL, MetaTool). The one JevBench row — 198 against Jev's 200 of 231 — sits in the bottom-right tile marked 'gap not significant,' a two-item loss. (MaestroYan/JEMM model card, 'Accuracy.')"
/>

The figure is admirably honest and the post is not. The six wins JEMM charts big
are on its **own held-out sets**: web actions on Mind2Web (36.0 to 54.6 with a
screenshot, Jev given only text; n=894), tool calling on BFCL (88.7 to 93.3;
n=2,148), tool abstention on MetaTool (92.6 to 99.3). Those are real, paired,
with lower bounds. But the one row on an outside benchmark, **JevBench public, is
198 of 231 against Jev's 200** (measured, from the figure) — a loss, filed under
"on par, or gap not significant." That is 85.7%, the post's headline number —
above Laya's 58.4%, below Jev's. There is no sealed-tier number, and no
calibration number anywhere.

The latency figure is as careful and has the same shape: single-question medians
of 271, 274 and 299 ms against Jev's 455, 454 and 458 — 0.60× to 0.65×, a −31.2%
aggregate across the five text configs (measured, from 600 paired requests). But
JEMM runs on your GPU and Jev is a hosted API, so "0.60× of Jev's time" is partly
the network round-trip to TypeSafe, not compute — local-versus-hosted is the most
common unstated confound in this category's latency claims.

### Lumma-fev (FrontiersMind): the training write-up, the same copied column

[Week two](/articles/jev-alternatives-week-two) covered Lumma-fev's 0.1B and 0.6B
and noted the 4B and 9B were promised but not on Hugging Face. They are now, and
the post this week is the [training write-up](https://www.frontiersmind.ai/blog/lumma-fev/index.html).

Measured, from the safetensors metadata and `config.json`: the 4B is
4,207,062,528 parameters with a **1,311,232-parameter FP32 head** — exactly
Kev-4B's pointer-head size, the same count [NeoHorse-Jev-4B](/articles/jev-alternatives-week-two)
borrowed — and the 9B is 7,938,782,208 with a 2,097,664-parameter head. The
config sets `option_isolation: false`, so it is a pointer: options attend to each
other inside a question, and order matters. The blog adds the training recipe,
which is the genuinely new content: the **4B is continual pre-training on the
Qwen3.5-4B line, then decision fine-tuning**; the 9B the same on the 9B line; the
0.6B is a frozen backbone with a LoRA; and the 0.1B is a from-scratch
Nandi-Mini-150M backbone, fully fine-tuned.

What the blog does *not* add is a re-measured Jev. The card's four-benchmark
table (measured, arithmetic recomputed) puts the 4B at an average of 0.90 and Jev
at 0.75 across Typed-decisions, AG News, DAIR Emotion and Banking77, with
[GLiNER-2.5-Decide](/articles/gliner-2-5-decide) fourth at 0.54 — but the Jev
column (0.72, 0.91, 0.48, 0.87) is identical, to the digit, to the figures on
Laya's own model card that week two already flagged, including DAIR Emotion's
implausibly low 0.48 for a frontier decision model. Only one of the four,
Typed-decisions, is decision-native, and there the 4B leads 0.78 to 0.72. No n,
no eval code, no calibration. See [Laya vs Jev](/articles/laya-vs-jev) for the
background on the copied column.

## The serving layer: SGLang makes the readout first-class

The most consequential post of the week is not a model. SGLang added
`/v1/decisions` and `/v1/systemone`, and that is the [any model can be
Jev](/articles/any-model-can-be-jev) thesis shipped as product.

That piece found the System One readout was a serving feature fifteen months older
than the category: one prefill, `max_new_tokens: 0`, read the logprobs of the
tokens you name. `/v1/decisions` (measured, from the docs) is that, wrapped: you
send an input and typed questions, and "each answer comes back with the
probability of every option, read from the model's next-token scores at the answer
position. No text is generated and no output is parsed. It needs no special
checkpoint." Choice questions take 2 to 26 options labelled `A`–`Z`, score
questions 2 to 10 levels labelled `0`–`9`, one prefill each — a letter readout,
with all its limits.

Two details are better than most of the models'. The response carries
**`label_mass`**, "the full-vocabulary probability of the answer labels at the
answer position. A low value means the model puts most of its probability outside
the offered answers" — exactly the abstention signal the any-model piece said
renormalisation throws away, surfaced as a first-class field. And the docs are
blunt: "None of these values is a calibrated probability that the decision is
correct. Validate any threshold on labeled data from your workload." What they do
not touch is the batching non-determinism [Jev is not
deterministic](/articles/jev-is-not-deterministic) found — a scored answer read
off the logits can still move between batch sizes, serving layer or not.

`/v1/systemone` is the compatibility layer: it "serves the same decisions in the
request and response shape of the System One API," so "the official TypeSafe
SDKs" work by pointing their base URL at the server. It reaches Jev's full 255
options, but only by handing every option past 26 a two-letter label (`AA`, `AB`,
…) that the docs warn "include common words and have unequal priors, so answers
above 26 options can depend on option order." The demo attached to the post —
Qwen3.8-27B turned into a multimodal player that cleared Pokémon FireRed's Elite
Four with sub-100 ms decisions — is a reported showcase, not in the docs. This is
the infrastructure that makes the models above interchangeable: one API, any
backbone, no fine-tune — and no calibration guarantee either.

## Packaging: Laya on Unsloth

Unsloth's post is a deployment story, not a model: "run Laya Decision models
locally on just 4GB RAM," on CPU, Mac, Windows, Linux or GPU, "through a
Jev-compatible API via Unsloth Desktop." The docs (reported) give the real shape:
the default multilingual Laya is a **678 MB** model needing 4 GB of RAM with a
1,024-token context, the English and typed-decision variants are 846 MB; CPU is
the default, the GPU is optional and falls back to CPU on an out-of-memory. The
Unsloth GUI "expose[s] Laya through a TypeSafe-compatible Jev API, so existing Jev
integrations can work by pointing them at your local server."

There is nothing new to benchmark here — Laya is the same model
[Laya vs Jev](/articles/laya-vs-jev) and [Laya on MLX](/articles/laya-mlx)
covered — and that is the point. The "4 GB RAM" is the host requirement, not the
model; a 678 MB encoder classifier was already small. What Unsloth adds is the
one-click local server speaking Jev's wire format, which is the same thing
localjev, lumma-fev-serve and now SGLang each built independently. The wire format
is by now reproduced about as many times as the readout.

## The oddball: taiga-s1 builds parts in FreeCAD

`shhivv/taiga-s1` is the one that does not fit the frame, and it is the most fun.
A **1,228,163-parameter** model (measured, safetensors total — "1.2M"), trained
from scratch, no language model and no vision model, that builds 3D parts in
FreeCAD. "Your agent plans, taiga executes": a planner hands it an ordered
feature list ("plate 40×30×10, Ø6 hole, polar pattern ×6, fillet the top edges")
and taiga picks the next FreeCAD command, step by step, from the ones currently
available, at **about 1 ms per decision on a CPU** (reported).

It is a decision model in the strict sense this series uses: a 3-layer encoder
and 2-layer decoder (measured, `config.json`: `S1Model`, width 128, 4 heads)
where "candidate commands attend to the state and the active goal item, and each
gets one score" — a per-option scorer over the available actions, with a fitted
temperature and calibrated probabilities. And the calibration is the only one
this week I could read straight from the repo: `config.json` records a held-out
set of 5,573 states at **ECE 0.0272** after temperature scaling (0.0416 at T=1),
accuracy 0.9551, fitted temperature 2.554 (measured).

It makes no claim against Jev and belongs on no JevBench tier. It is here because
it is the clearest answer to "what is a decision model, minimally?" — a tiny head
that scores a constrained set of typed options and knows how confident it is. You
do not need 27 billion parameters for that. You need 1.2 million, if the option
set is a CAD workbench rather than the open world.

## The evaluation went backwards

Line up what week three measured against what week two did, and the trend is not
the one the posts imply.

**Nobody is on the sealed board.** Week two put XOR and Open-Jev on JevBench's 308
sealed items, where both sat below Jev. Week three's four new models are on *no*
sealed tier: clef and Lumma publish their own suites, Intern-Decision runs its own
copy of the public tiers, and JEMM's only JevBench number is a public one it
loses. The sealed items were the whole point of week two's finding — a public
lead of a few points said nothing about the sealed ranking, and
[reproducing Jev](/articles/reproducing-jev) showed how quickly a board fills with
rows aimed at the public items — and this week not one new release went near them.

**The baseline is not a constant.** Jev is a single closed model, and this week's
publishers report its median single-request latency as 106.30 ms (Intern-Decision,
on a 4090 "local HF"), 256 ms (Lumma's card), 455 ms (JEMM's paired run) and 524.1
ms (clef's Decision Index). That is a **5× spread on the same model**, because each
ran it a different way — some against the hosted API with its network hop, some in
a local configuration I cannot reconstruct for a closed model. A "−31.2% vs Jev"
is a statement about the publisher's stopwatch as much as about the model.

**Calibration is still the tell.** The one property you cannot fake from the
serving layer is the one that separates the careful releases. taiga-s1 ships a
measured ECE in its config. Intern-Decision reports ECE and Brier and a
calibration pilot. clef reports a ForecastBench Brier. JEMM and Lumma-fev publish
none, and SGLang and Unsloth say plainly that what they return is not calibrated
and you must measure your own. That split — who publishes a reliability number —
is still, as it was in [any model can be Jev](/articles/any-model-can-be-jev), the
fastest way to tell who did the expensive half.

## If you are choosing this week

All reasoned, from the evidence above:

- **A real trained head, documented:** clef — the first joint scorer trained as a
  unit, with the fullest card. Every number on it is Cloudflare's own run, on no
  third-party board.
- **A calibration temperature that ships:** Intern-Decision and taiga-s1 both fit
  one and tell you the value; taiga's held-out ECE is in the repo.
- **More than 26 options:** Lumma-fev's pointer (255) or Intern-Decision's 62
  symbols. On SGLang you reach 255 only through order-dependent two-letter labels.
- **Drop-in on your own backbone:** SGLang `/v1/systemone`, no fine-tune — and no
  calibration guarantee.
- **A sealed-item or independent JevBench score:** nothing this week. The week-two
  releases on the board are still the only ones with one.

## The take

Week three is the strongest set of open decision models yet — a trained joint
head, a 62-symbol readout, a serving layer that makes the whole thing a flag, a
1.2-million-parameter CAD decider — and the weakest set of evaluations. Every
"beats Jev" number is the publisher's own, most are not on JevBench at all, none
is sealed, and the Jev they are beating has a latency that depends on who held the
stopwatch.

The one release that measured a reliability number straight into its config file
is the 1.2-million-parameter one that does not claim to beat anything. That is the
week, in one sentence.

---

*Sources, read on 2026-10-02: [Cloudflare/clef](https://huggingface.co/Cloudflare/clef) and [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash); [internlm/Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B), [-2B](https://huggingface.co/internlm/Intern-Decision-2B) and [-0.8B](https://huggingface.co/internlm/Intern-Decision-0.8B); [MaestroYan/JEMM](https://huggingface.co/MaestroYan/JEMM); the [Lumma-fev collection](https://huggingface.co/collections/FrontiersMind/lumma-decision-models) and its [training write-up](https://www.frontiersmind.ai/blog/lumma-fev/index.html); [shhivv/taiga-s1](https://huggingface.co/shhivv/taiga-s1) and [its code](https://github.com/shhivv/taiga-s1); SGLang's [decision-models docs](https://docs.sglang.io/docs/supported-models/decision_models); and Unsloth's [Laya + Jev API guide](https://unsloth.ai/docs/models/decision-laya). Parameter counts and head architectures are safetensors header and metadata reads; no third-party code was executed. The figure is reproduced for commentary from JEMM's model card (Apache-2.0); the interactive is original.*
