~/satyajit

Jev alternatives, week three: bigger models, smaller benchmarks

mdjsonmcp

2026-10-02 · 18 min · explainer · llm · ecosystem · benchmarks · calibration · architecture

Week one was code with almost no numbers. Week two brought seven open-weight decision models with numbers, and the finding was that most of the comparisons stood on JevBench's 231 public items, that on the 308 sealed items a three-point public lead over Jev turned into a 2.9-point deficit, and that two releases out of twelve measured Jev themselves and printed the loss.

Week three is seven more posts, and the releases are, by some distance, the most serious in the category so far. The architectures stopped being merged LoRAs with the vision tower left untouched: this week has a trained 128-million-parameter cross-attention head from Cloudflare, a readout over 62 symbols from Shanghai AI Lab, a serving layer that makes the whole trick a flag, and a from-scratch 1.2-million-parameter model that does not speak text at all.

The evaluation went the other way. Not one of this week's new models carries an independent or sealed-item JevBench score. clef and Lumma-fev publish their own suites and never touch it; Intern-Decision runs its own copy of the public tiers; and JEMM's lone JevBench row is a loss it files under "gap not significant." Week two put two releases on the sealed board. Week three: none.

Parameter counts below are safetensors header or metadata sums, not the cards' round numbers. Code and configs were read, never run. Measured means read from a file or recomputed from published rows; reported means the author's claim, which I could not check; reasoned means my inference. All checked on 2026-10-02.

The one idea, again

A typed decision model does one thing a chat model behind a harness does not: it answers a schema of typed questions — choice, score, yes/no — by returning a calibrated probability for every option in a single forward pass, with no text generated and nothing to parse. A chat model writes a label into JSON and you hope the format holds; a decision model reads the answer off the logits, or off a trained head, at the position where the answer goes.

Week two sorted the open ones into three families by how they turn options into probabilities, and the same split holds this week:

From each release's card, config and head tensors (measured): Intern-Decision and JEMM are letter readouts. Lumma-fev is a pointer. clef is the new one — a trained joint scorer. taiga-s1 is a per-command scorer for a domain that has nothing to do with language. SGLang is not a model; it is the serving layer that makes a letter readout out of anything. Laya on Unsloth is a packaging of an existing model.

week three · seven entrants, and who ran the Jev comparison

Jev's median single-request latency, as four publishers measured it (ms, log axis)

50
100
200
500
106.3
256
455
524.1
  • 106.3 ms — Intern-Decision (RTX 4090, local HF)
  • 256 ms — Lumma-fev (card P50)
  • 455 ms — JEMM (paired, short)
  • 524.1 ms — clef (Decision Index median)

Intern-Decision Shanghai AI Lab

new model
size
0.85B / 2.2B / 4.5B params
modality
text + up to 8 images
decides by
symbol readout (A–Z, a–z, 0–9)
option cap
62 options, 16 questions
vs Jev
4B avg 90.02 vs Jev 88.74 across seven benches (self-run)
calibration
4B ECE 0.065 vs Jev 0.095 (self-run)
license
Apache-2.0
evidence: self-run · JevBench public

JEMM MaestroYan

new model
size
117M LoRA on Qwen3.8-27B
modality
text + screenshot
decides by
letter readout (A–Z, 0–5)
option cap
32 options
vs Jev
wins its own held-out splits; JevBench public 198 vs 200 of 231 (−2)
calibration
none published
license
Apache-2.0
evidence: self-run · JevBench public

Lumma-fev FrontiersMind

new model
size
0.15B / 0.6B / 4B / 9B params
modality
text
decides by
Kev-style pointer head
option cap
255 options
vs Jev
4B avg 0.90 vs Jev 0.75 on four tasks; Jev column copied from Laya's card
calibration
none published
license
Apache-2.0
evidence: self-run · own held-out set

clef / clef-flash Cloudflare

new model
size
27.4B + 128M head / 9.4B + 122M head
modality
text, JSON, images, video
decides by
trained joint schema head
option cap
schema-scored
vs Jev
wins most of its own Decision Index; Jev wins the reasoning benches; no JevBench
calibration
ForecastBench Brier 13.9 vs Jev 17.4 (self-run)
license
Apache-2.0
evidence: no JevBench number at all

SGLang /v1/decisions SGLang

serving layer
size
any served chat model
modality
text (VLM in the demo)
decides by
label logprobs via /v1/score
option cap
26 (255 with two-letter labels)
vs Jev
/v1/systemone speaks the TypeSafe SDK; no accuracy claim in the docs
calibration
label_mass only; docs say it is not calibrated
license
Apache-2.0 (SGLang)
evidence: makes no beats-Jev claim

Laya on Unsloth Unsloth

packaging
size
678–846 MB (the Laya model)
modality
text
decides by
Laya, unchanged
option cap
Laya's
vs Jev
runs Laya locally behind a Jev-compatible API; no new numbers
calibration
Laya's own
license
Laya + Unsloth
evidence: makes no beats-Jev claim

taiga-s1 shhivv

oddball
size
1.2M params, from scratch
modality
FreeCAD state (no text model)
decides by
per-command scorer
option cap
available commands
vs Jev
not a Jev competitor: a ~1 ms FreeCAD action decider
calibration
held-out ECE 0.0272 (measured, in config)
license
MIT
evidence: makes no beats-Jev claim
No entrant this week carries an independent or sealed-item JevBench score. The chip colour is who ran the comparison, not how it turned out.

The new models

clef (Cloudflare): the first one that trained a real head

The post is a one-liner — "another open source Jev, this time by Cloudflare, post-trained on Qwen 3.8" — and it undersells the release. Cloudflare/clef is the most complete thing in the category so far.

Cloudflare/clef@2f3de3d · snapshot 2026-10-02
parameters
27.36B
repo size
54.99 GB
architecture
Qwen3_5ForConditionalGeneration
task
image-text-to-text
library
transformers
license
apache-2.0
safetensors
13 shards
largest file
4.99 GB
files
25
downloads
18
likes
458
parameters by dtype
BF1627.36B
clefcloudflaresystemoneqwen3.8post-trainimage-text-to-typed-outputmultimodalstructured-output

A full 27,356,728,560-parameter Qwen3.8-27B backbone (BF16, twelve shards, vision encoder included) plus a separate joint_head.safetensors of 128,056,324 parameters, counted from the safetensors headers by range request. clef-flash is a 9,409,813,744-parameter backbone with a 121,762,820-parameter head. Apache-2.0, following the base model.

repo last modified 2026-10-01

What makes it different is the head, and I read it tensor by tensor (measured, joint_head.safetensors, joint_head_config.json). It is not a letter readout and not a LoRA. It is a 128-million-parameter cross-attention transformer: two "evidence" layers that route information from the state to each question, then four transformer layers, projections pulling the backbone's 5,120-wide hidden states down to a 1,024-wide working space, and a residual scorer MLP. The card's own words: it "scores all options of all questions jointly." So unlike every isolated scorer before it, clef's options do see each other — the first head in this category trained as a unit rather than folded in as an adapter. It also reads images and video, where Jev-Omni first pushed the category multimodal.

The benchmark table is the biggest of the week — about 42 rows — and it is also the catch. Every number is from "our internal run of the Decision Index 0.2.1 suite" (reported), Cloudflare's own benchmark. clef wins most of it: BANKING77 macro-F1 94.2 against Jev's 79.7, CLINC150 97.4 against 89.3, ForecastBench Brier 13.9 against Jev's 17.4 (lower is better). But read the rows where Jev wins, because they are the reasoning-heavy ones: GPQA Diamond 78.3 against clef's 48.0, MMLU-Pro 82.7 against 65.9, BBH 92.9 against 73.7. The pattern (reasoned) is clean: clef wins routing, classification and structured extraction; Jev wins the benches that need actual inference. Median latency on the same suite: clef 209.3 ms, clef-flash 38.8 ms, Jev 524.1 ms (reported).

None of this is JevBench. clef is absent from the one third-party board the rest of the category reports on, so there is no way to rank it against the week-two releases: the best-documented card of the wave is the one you can place on no independent axis.

Intern-Decision (Shanghai AI Lab): 62 symbols, and a fitted temperature

The post, from ModelScope, is specific: three sizes averaging 79.38, 84.68 and 90.02 across seven decision benchmarks, the 4B "surpasses Jev at 88.74 while achieving better probability calibration," and mean latency **33.98, 33.28 and 44.16 ms against 109.70 ms for Jev in the same local HF setup."

internlm/Intern-Decision-4B@0e5e6aa · snapshot 2026-10-02
parameters
4.54B
repo size
9.10 GB
finetuneQwen/Qwen3.5-4B
architecture
Qwen3_5ForConditionalGeneration
task
image-text-to-text
library
transformers
license
apache-2.0
safetensors
4 shards
largest file
4.26 GB
files
21
downloads
1.3K
likes
73
parameters by dtype
BF164.54BF323.8K
decision-makingmultimodalstructured-prediction

Base Qwen3.5-4B. Safetensors totals (measured): 0.8B is 852,985,920, the 2B 2,213,241,664 and the 4B 4,539,265,536 parameters. Apache-2.0, with the Qwen3.5 upstream notices retained. A separately fitted candidate-probability temperature of 1.99241824 ships with the 4B.

repo last modified 2026-09-26

The mechanism is a letter readout with two refinements (measured, from the card's inference section). It maps each question's options to single-token symbols — A–Z, then a–z, then 0–9, 62 of them — so the cap is 62 options where XOR and JEMM stop at 26. It renders a full JSON answer skeleton with one <decision> placeholder per field, runs one forward pass, and reads the logits at the position immediately before each placeholder, so up to 16 questions are answered in that single pass. Then it applies a fitted candidate-probability temperature — softmax(log(p) / T) with T = 1.99241824, tuned by NLL minimisation on 1,728 held-out cases — which "updates confidence while preserving the argmax decision." It is still a vocabulary readout, so A and B carry different priors and order still matters; the card does not mention reversing and averaging.

The arithmetic checks (measured): the 4B's seven scores are 100.00, 98.61, 73.87, 80.55, 96.45, 90.82 and 89.86, which average to 90.02, and Jev's seven average to 88.74. So "surpasses Jev at 88.74" means the 4B's 90.02 beats Jev's 88.74 — the number in the headline is Jev's. On calibration the 4B reports ECE 0.065 against Jev's 0.095 and Brier 0.347 against 0.358, and a separate 96-case pilot puts the 4B at ECE 0.089 after calibration against Jev's 0.130 (all reported).

Two catches. Three of the seven benchmarks are JevBench-Easy, -Original and -Hard — the public tiers, the 48 + 72 + 111 that make up the 231 public items — run by internlm's own harness, not JevBench's. And the Jev comparison, accuracy and latency both, is internlm running Jev itself: the latency table says "Jev in the same local HF setup" at 109.70 ms, but Jev is TypeSafe's hosted, closed model, and the card never explains how it ran locally on a 4090. I could not verify that number, or that the Jev benchmark rows are the hosted model's.

JEMM (MaestroYan): the candid figure, and the loud post

The post, in Chinese, leads with "85.7% crushes Laya, tool-calling 93.3% overtakes Jev." MaestroYan/JEMM is a LoRA adapter (measured, adapter_config.json): r=16, α=32, 116,727,808 parameters, no head tensor at all, on Qwen/Qwen3.8-27B. It is a pure letter readout — labels A–Z and 0–5, 2 to 32 candidates — that reads the base model's own last-token logits over those label tokens and softmaxes them. Multimodal, because Qwen3.8 is: it takes a screenshot (trained at 1280×800). Its training data is listed — Mind2Web, BFCL, ToolACE, Banking77, CLINC150 and more — and "no Jev outputs were used."

A lollipop chart titled 'JEMM vs Jev 1.13 — Accuracy, same sealed questions, paired one by one.' Six rows show JEMM ahead of Jev 1.13: web actions with screenshot 36.0 to 54.6 (+18.6 pp), web actions text only 35.2 to 45.8 (+10.7), content moderation category 54.4 to 63.2 (+8.8), tool abstention 92.6 to 99.3 (+6.7), tool calling 88.7 to 93.3 (+4.6), content moderation 90.2 to 93.0 (+2.8). A lower 'Also measured — on par, or gap not significant' panel has six tiles: batched 4 questions 75 vs 65, batched 8 questions 51 vs 40, Banking77 intent 89.0 vs 87.0, CLINC150 intent 89.5 vs 88.8, MetaTool similar tools 71.2 vs 72.7 (−1.5), and JevBench public 198 vs 200 of 231 (−2).
JEMM's own accuracy figure. The six headline wins are on JEMM's own held-out splits (Mind2Web, BFCL, MetaTool). The one JevBench row — 198 against Jev's 200 of 231 — sits in the bottom-right tile marked 'gap not significant,' a two-item loss. (MaestroYan/JEMM model card, 'Accuracy.')

The figure is admirably honest and the post is not. The six wins JEMM charts big are on its own held-out sets: web actions on Mind2Web (36.0 to 54.6 with a screenshot, Jev given only text; n=894), tool calling on BFCL (88.7 to 93.3; n=2,148), tool abstention on MetaTool (92.6 to 99.3). Those are real, paired, with lower bounds. But the one row on an outside benchmark, JevBench public, is 198 of 231 against Jev's 200 (measured, from the figure) — a loss, filed under "on par, or gap not significant." That is 85.7%, the post's headline number — above Laya's 58.4%, below Jev's. There is no sealed-tier number, and no calibration number anywhere.

The latency figure is as careful and has the same shape: single-question medians of 271, 274 and 299 ms against Jev's 455, 454 and 458 — 0.60× to 0.65×, a −31.2% aggregate across the five text configs (measured, from 600 paired requests). But JEMM runs on your GPU and Jev is a hosted API, so "0.60× of Jev's time" is partly the network round-trip to TypeSafe, not compute — local-versus-hosted is the most common unstated confound in this category's latency claims.

Lumma-fev (FrontiersMind): the training write-up, the same copied column

Week two covered Lumma-fev's 0.1B and 0.6B and noted the 4B and 9B were promised but not on Hugging Face. They are now, and the post this week is the training write-up.

Measured, from the safetensors metadata and config.json: the 4B is 4,207,062,528 parameters with a 1,311,232-parameter FP32 head — exactly Kev-4B's pointer-head size, the same count NeoHorse-Jev-4B borrowed — and the 9B is 7,938,782,208 with a 2,097,664-parameter head. The config sets option_isolation: false, so it is a pointer: options attend to each other inside a question, and order matters. The blog adds the training recipe, which is the genuinely new content: the 4B is continual pre-training on the Qwen3.5-4B line, then decision fine-tuning; the 9B the same on the 9B line; the 0.6B is a frozen backbone with a LoRA; and the 0.1B is a from-scratch Nandi-Mini-150M backbone, fully fine-tuned.

What the blog does not add is a re-measured Jev. The card's four-benchmark table (measured, arithmetic recomputed) puts the 4B at an average of 0.90 and Jev at 0.75 across Typed-decisions, AG News, DAIR Emotion and Banking77, with GLiNER-2.5-Decide fourth at 0.54 — but the Jev column (0.72, 0.91, 0.48, 0.87) is identical, to the digit, to the figures on Laya's own model card that week two already flagged, including DAIR Emotion's implausibly low 0.48 for a frontier decision model. Only one of the four, Typed-decisions, is decision-native, and there the 4B leads 0.78 to 0.72. No n, no eval code, no calibration. See Laya vs Jev for the background on the copied column.

The serving layer: SGLang makes the readout first-class

The most consequential post of the week is not a model. SGLang added /v1/decisions and /v1/systemone, and that is the any model can be Jev thesis shipped as product.

That piece found the System One readout was a serving feature fifteen months older than the category: one prefill, max_new_tokens: 0, read the logprobs of the tokens you name. /v1/decisions (measured, from the docs) is that, wrapped: you send an input and typed questions, and "each answer comes back with the probability of every option, read from the model's next-token scores at the answer position. No text is generated and no output is parsed. It needs no special checkpoint." Choice questions take 2 to 26 options labelled A–Z, score questions 2 to 10 levels labelled 0–9, one prefill each — a letter readout, with all its limits.

Two details are better than most of the models'. The response carries label_mass, "the full-vocabulary probability of the answer labels at the answer position. A low value means the model puts most of its probability outside the offered answers" — exactly the abstention signal the any-model piece said renormalisation throws away, surfaced as a first-class field. And the docs are blunt: "None of these values is a calibrated probability that the decision is correct. Validate any threshold on labeled data from your workload." What they do not touch is the batching non-determinism Jev is not deterministic found — a scored answer read off the logits can still move between batch sizes, serving layer or not.

/v1/systemone is the compatibility layer: it "serves the same decisions in the request and response shape of the System One API," so "the official TypeSafe SDKs" work by pointing their base URL at the server. It reaches Jev's full 255 options, but only by handing every option past 26 a two-letter label (AA, AB, …) that the docs warn "include common words and have unequal priors, so answers above 26 options can depend on option order." The demo attached to the post — Qwen3.8-27B turned into a multimodal player that cleared Pokémon FireRed's Elite Four with sub-100 ms decisions — is a reported showcase, not in the docs. This is the infrastructure that makes the models above interchangeable: one API, any backbone, no fine-tune — and no calibration guarantee either.

Packaging: Laya on Unsloth

Unsloth's post is a deployment story, not a model: "run Laya Decision models locally on just 4GB RAM," on CPU, Mac, Windows, Linux or GPU, "through a Jev-compatible API via Unsloth Desktop." The docs (reported) give the real shape: the default multilingual Laya is a 678 MB model needing 4 GB of RAM with a 1,024-token context, the English and typed-decision variants are 846 MB; CPU is the default, the GPU is optional and falls back to CPU on an out-of-memory. The Unsloth GUI "expose[s] Laya through a TypeSafe-compatible Jev API, so existing Jev integrations can work by pointing them at your local server."

There is nothing new to benchmark here — Laya is the same model Laya vs Jev and Laya on MLX covered — and that is the point. The "4 GB RAM" is the host requirement, not the model; a 678 MB encoder classifier was already small. What Unsloth adds is the one-click local server speaking Jev's wire format, which is the same thing localjev, lumma-fev-serve and now SGLang each built independently. The wire format is by now reproduced about as many times as the readout.

The oddball: taiga-s1 builds parts in FreeCAD

shhivv/taiga-s1 is the one that does not fit the frame, and it is the most fun. A 1,228,163-parameter model (measured, safetensors total — "1.2M"), trained from scratch, no language model and no vision model, that builds 3D parts in FreeCAD. "Your agent plans, taiga executes": a planner hands it an ordered feature list ("plate 40×30×10, Ø6 hole, polar pattern ×6, fillet the top edges") and taiga picks the next FreeCAD command, step by step, from the ones currently available, at about 1 ms per decision on a CPU (reported).

It is a decision model in the strict sense this series uses: a 3-layer encoder and 2-layer decoder (measured, config.json: S1Model, width 128, 4 heads) where "candidate commands attend to the state and the active goal item, and each gets one score" — a per-option scorer over the available actions, with a fitted temperature and calibrated probabilities. And the calibration is the only one this week I could read straight from the repo: config.json records a held-out set of 5,573 states at ECE 0.0272 after temperature scaling (0.0416 at T=1), accuracy 0.9551, fitted temperature 2.554 (measured).

It makes no claim against Jev and belongs on no JevBench tier. It is here because it is the clearest answer to "what is a decision model, minimally?" — a tiny head that scores a constrained set of typed options and knows how confident it is. You do not need 27 billion parameters for that. You need 1.2 million, if the option set is a CAD workbench rather than the open world.

The evaluation went backwards

Line up what week three measured against what week two did, and the trend is not the one the posts imply.

Nobody is on the sealed board. Week two put XOR and Open-Jev on JevBench's 308 sealed items, where both sat below Jev. Week three's four new models are on no sealed tier: clef and Lumma publish their own suites, Intern-Decision runs its own copy of the public tiers, and JEMM's only JevBench number is a public one it loses. The sealed items were the whole point of week two's finding — a public lead of a few points said nothing about the sealed ranking, and reproducing Jev showed how quickly a board fills with rows aimed at the public items — and this week not one new release went near them.

The baseline is not a constant. Jev is a single closed model, and this week's publishers report its median single-request latency as 106.30 ms (Intern-Decision, on a 4090 "local HF"), 256 ms (Lumma's card), 455 ms (JEMM's paired run) and 524.1 ms (clef's Decision Index). That is a 5× spread on the same model, because each ran it a different way — some against the hosted API with its network hop, some in a local configuration I cannot reconstruct for a closed model. A "−31.2% vs Jev" is a statement about the publisher's stopwatch as much as about the model.

Calibration is still the tell. The one property you cannot fake from the serving layer is the one that separates the careful releases. taiga-s1 ships a measured ECE in its config. Intern-Decision reports ECE and Brier and a calibration pilot. clef reports a ForecastBench Brier. JEMM and Lumma-fev publish none, and SGLang and Unsloth say plainly that what they return is not calibrated and you must measure your own. That split — who publishes a reliability number — is still, as it was in any model can be Jev, the fastest way to tell who did the expensive half.

If you are choosing this week

All reasoned, from the evidence above:

The take

Week three is the strongest set of open decision models yet — a trained joint head, a 62-symbol readout, a serving layer that makes the whole thing a flag, a 1.2-million-parameter CAD decider — and the weakest set of evaluations. Every "beats Jev" number is the publisher's own, most are not on JevBench at all, none is sealed, and the Jev they are beating has a latency that depends on who held the stopwatch.

The one release that measured a reliability number straight into its config file is the 1.2-million-parameter one that does not claim to beat anything. That is the week, in one sentence.


Sources, read on 2026-10-02: Cloudflare/clef and Cloudflare/clef-flash; internlm/Intern-Decision-4B, -2B and -0.8B; MaestroYan/JEMM; the Lumma-fev collection and its training write-up; shhivv/taiga-s1 and its code; SGLang's decision-models docs; and Unsloth's Laya + Jev API guide. Parameter counts and head architectures are safetensors header and metadata reads; no third-party code was executed. The figure is reproduced for commentary from JEMM's model card (Apache-2.0); the interactive is original.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Jev alternatives, week three: bigger models, smaller benchmarks", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026jevalternativesweekthree,
  author = {Satyajit Ghana},
  title  = {Jev alternatives, week three: bigger models, smaller benchmarks},
  url    = {https://ai.thesatyajit.com/articles/jev-alternatives-week-three},
  year   = {2026}
}
share