2026-10-06 · 23 min · explainer · llm · ecosystem · benchmarks · calibration · multimodal · on-device · inference
Week two was seven open decision models and the discovery that most "beats Jev" numbers stood on JevBench's public items. Week three was the most serious architectures yet and the thinnest evaluations: not one new model on a sealed or independent board.
Week four is four posts, and none is mainly about a better model. All four
are about where a decision model runs: behind a closed API that now accepts images (Liquid's
d1), as open weights in six sizes loaded through Transformers (vLLM Semantic
Router's Decision 2.0), natively on a Mac through Core ML (FluidUse), and
inside llama.cpp's server behind the same /v1/systemone endpoint as Jev
itself.
The evaluation story inverts. The one publisher that shipped its raw results is the closed one. The open family leaves Jev out of its tables; the on-device port checks that it matches upstream, never whether it is right; llama.cpp makes no accuracy claim and tells you to calibrate your own cutoffs.
Measured means I read it from a file, a safetensors header or source code, or recomputed it from a published file; reported means the publisher's figure, which I could not re-run; reasoned means my inference from the other two. Code was read, never executed. All checked on 2026-10-06.
The one idea, again
A decision model answers a schema of typed questions — choice, score,
noul (yes/no) — by returning a probability for every option, read in one
forward pass, with no tokens generated and nothing to parse. The full
treatment is in what Jev actually is, why
any model can be Jev and what decision models
cannot do; and Jev is not
deterministic is why "one pass" does not
mean "same answer every time."
The open ones still sort into three readout families: a letter readout (options become letters, read the letters' next-token logits), a pointer (each option's end-token hidden state is scored against the question's, as in Kev), and a joint scorer (a trained head scores everything at once, as in clef). This week (measured, from configs and code): d1 is closed and undisclosed; Decision 2.0 is a pointer with one head design across six sizes; FluidUse re-implements that pointer in Core ML; and llama.cpp is not a model but six readouts behind one endpoint.
d1 with vision (Liquid): the best receipt of the week is a JSON file
Liquid launched d1 as a text model on 29 September. The week-four post adds images: "images, text or both as inputs," tested "against GPT-6.1 Sol and Claude Opus 5.5 on six real applications," matching or beating GPT-6.1 Sol "on four of them," at "19x to 200x less" cost (reported).

What can be checked is the interface. The request goes to
https://api.liquid.ai/decisions/v1/systemone, the docs install TypeSafe's own
typesafe-sdk and point its base_url at Liquid, and the question types are
Jev's noul, choice and score — so d1 is a drop-in for Jev's wire format,
the same thing week three watched
SGLang and Unsloth build. Billing is input tokens only, at $0.04 per million.
An image costs 1.5 tokens per 32×32 patch, so a 1024×1024 image is 1,536
tokens, and "every question is charged for its text and all the images again"
(measured, from the docs). That last clause matters for vision: ask five
questions about one photo and you pay for the photo five times. Requests take
up to 8 images and 10,000 patches.
The chart, and the file behind it

The d1 playground is a static site, and its comparison panel loads
demos/data/compare.json: per application and per model, quality, cost,
seconds, request counts and token counts, stamped "measured 2026-10-05". I
recomputed the chart from it (measured):
- Quality averages to 93.6%, 91.3% and 98.2% — the chart's 94, 91 and 98.
- Time averages to 3.74, 14.15 and 15.75 seconds — the chart's 3.7, 14.1 and 15.8.
- Cost reproduces only one way. With Smart Folders' cost taken per passage (the file stores it per 105 passages) and every other application per run, the means are $0.54, $25.5 and $84.7 per 1,000 runs — the chart's $0.54, $26 and $85. The blog's methodology confirms that Smart Folders' "cost is per 1,000 passages."
So the chart is honest arithmetic over an odd unit: "cost per 1,000 runs" averages a coding session, a passage, a 150-ticket query, a photo, a question and a web goal, and Code Search and Smart Filter make up most of the frontier models' averages (reasoned). The per-application cost ratios are the cleaner number: GPT-6.1 Sol costs 18.7x d1's on context compaction and up to 56.5x on Smart Folders; Opus costs 61.4x on compaction and 200.1x on Code Search (measured, from the file). That is exactly the post's "19x to 200x."
- one run
- cost: a mean over six different kinds of run; time: six different kinds of run
- quality
- mean of six different metrics
- sample
- six applications
quality (axis from 75%)
cost per 1,000 runs (log axis)
seconds per run
d1 is 47.2x cheaper than GPT-6.1 Sol and 156.5x cheaper than Claude Opus 5.5 here. Across single applications the range is 18.7x to 200.1x.
"Matches or beats GPT-6.1 Sol on four"
Per application (measured, from the file):
| application | d1 | GPT-6.1 Sol | Claude Opus 5.5 | sample |
|---|---|---|---|---|
| Context Compaction | 7 of 7 | 6 of 7 | 7 of 7 | 7 needed outputs |
| Visual Inspection | 90.6% | 82.3% | 91.7% | 96 photos |
| Web Agent | 7 of 7 | 7 of 7 | 7 of 7 | 7 goals |
| Smart Filter (F1) | 94.7 | 95.2 | 97.6 | 150 tickets |
| Smart Folders | 96.2% | 98.1% | 100% | 105 passages |
| Code Search | 12 of 15 | 13 of 15 | 15 of 15 | 15 questions |
d1 beats GPT-6.1 Sol on two (compaction, inspection), ties on one (web), and loses on three. The fourth "match" is Smart Filter, where the file's own display string rounds both F1 scores to "95%" — 94.7 against 95.2 unrounded. That is within any reasonable noise for 150 tickets, so "matches" is defensible; it is also the generous reading. Against Opus, d1 ties on compaction and web and trails on the other four.
Each application was run once, on seven to 150 items. Liquid's methodology note adds something most launch posts never do: "We wrote six of the 15 code questions and two of the four compaction sessions after d1's pipeline was set." That is a disclosure in d1's favour — those items were not tuned against — and a reminder of how small the sets are. The chat models answered in batches (48 requests for Smart Filter against d1's 900) at list prices without prompt caching. The chart's 3.7 seconds is per run with "up to 8 requests in flight"; the post's "text decisions in 200 to 300 ms" is a separate claim this file cannot check (reported).
The inspection result, and the threshold it needs
"85% to 97% accuracy across four VisA inspection tasks… without task-specific
training" (reported). The weights were not trained on VisA. The decision cut
was fitted per line. The playground's inspection/index.json stores, for each
production line, a threshold "calibrated at … on … good and … defective parts
that are not shown here" (measured, from the file and the page code):
| line | threshold on noul | calibration parts | balanced accuracy at calibration |
|---|---|---|---|
| circuit boards | 0.484 | 98 good, 88 defective | 0.952 |
| candles | 0.453 | 98 good, 87 defective | 0.902 |
| cashews | 0.126 | 48 good, 88 defective | 0.900 |
| chewing gum | 0.725 | 49 good, 88 defective | 0.960 |
Two things follow (reasoned). First, a noul of 0.5 does not mean the same
thing on every line: the cashew line calls a part defective at 0.126, the gum
line only at 0.725. Liquid's docs say the model "returns calibrated
probabilities"; a per-line cut ranging from 0.126 to 0.725 says that, for
defect detection at least, you still fit your own threshold on labelled
examples, and the demo did. Second, every model in the comparison also sees "a
good part from the same line next to the part to inspect," so this is
reference-based inspection, not zero-shot anomaly detection. The 85–97% range
itself I could not reproduce from these files: the calibration accuracies run
90–96%, and the comparison file's d1 figure over all 96 photos is 90.6%.
The games are reported and uncheckable without the harness: Tetris from 70 to 81 cleared lines when the screen is added to a text description, Wordle 12 of 12 solved in 3.8 guesses from screenshots, and Quick, Draw! at 5.2 of 6 doodles named among 62 words (random: 0.6).

The blog calls d1 "the first to extend these capabilities to images." It is
not: Jev-Omni went
multimodal first, clef and JEMM read images in week three, and a reply to the
launch post points at Mapika/decider-2b-vision, an Apache-2.0 vision decision
model on Hugging Face since 16 September (measured, from the Hub API). Its card
reports Visual7W 0.89. The one comparison to Jev, the chart above, is Liquid's own
run, and the release has no calibration number. Open weights are promised only
for "upcoming models."
Decision 2.0 (vLLM Semantic Router): six sizes, one head
The post is one line: "our newest state-of-the-art decision models, open in every size from 0.6B to 27B." The collection has six, each named, each Apache-2.0:
| model | base | backbone parameters (measured) | head parameters (measured) |
|---|---|---|---|
| Kai-0.6B | Qwen3-0.6B-Base | 596,049,920 | 1,053,184 |
| Eos-0.8B | Decision 1.0 Eos | 752,393,024 | 1,053,184 |
| Sol-2B | Decision 1.0 Sol | 1,881,825,088 | 2,105,856 |
| Nox-4B | Qwen3.5-4B-Base | 4,205,751,296 | 2,632,192 |
| Lux-9B | Decision 1.0 Lux | 7,936,684,544 | 4,211,200 |
| Vega-27B | Qwen3.8-27B + LoRA | rank-512 LoRA, 3,735,289,856 | 5,263,872 |
Backbone counts are safetensors index and header sums; head counts are the
decision_head.safetensors headers, read by range request.
- architecture
- Decision2Model
- task
- feature-extraction
- library
- transformers
- license
- apache-2.0
- safetensors
- 2 shards
- largest file
- 1.50 GB
- files
- 36
- downloads
- 725
- likes
- 29
Every weight of Qwen3-0.6B-Base fine-tuned plus a new 1,053,184-parameter FP32 endpoint head; the release is a uniform average of six fine-tunes (two recipes, three seeds). Raw probabilities at temperature 1, no calibration file. Apache-2.0.
repo last modified 2026-10-03
The head, from the code
The repositories ship their runtime as remote code, and
decision2/_vendor/dev2model/decision_model.py is short enough to read whole
(measured). Each question becomes its own prompt:
Context:
<state>
Task type: choice
Question:
<instructions>
Options:
<option>
{"key": "returns", "description": "Refunds, replacements and damaged deliveries"}
</option>
...
Select the single option best supported by the context and instructions.
Decision:The model runs that through the backbone once. It keeps the last hidden state
of each </option> block (the "candidate endpoints") and the hidden state of
the final token (the "global query"). The CandidateHead scores each option as
a bilinear term plus a small MLP term:
where and are the layer-normed option and query states and every projection is 256 wide. A softmax over the is the answer. The head is identical across sizes; only its input width changes (1,024 for Kai up to 5,120 for Vega), which is why it grows from about one to about five million parameters.
This is the pointer family. Because the backbone is causal, an option sees the options listed before it and not those after, so option order can still move the answer; there is no reverse-and-average in the runtime. The cap is 255 options. Every question is a separate prompt, but the questions share their context prefix, which is what lets the Core ML port below run them together.
Two more things the files say (measured). config.json carries "calibration": null, and the manifest calls the output "raw probabilities (temperature 1; no
calibration file)," with fixed offsets for five-level score questions; the
loss has an optional Brier term, so calibration is something training may push
toward, not something fitted after. And the releases are soups: Kai averages six
full fine-tunes, and Vega averages LoRA adapters with their "factors
concatenated along the rank," which is how a rank-512 adapter holds 3.7 billion
parameters.
The 27B's card says "29.37B" parameters. That is the Qwen3.8-27B text model's 25,624,600,064 plus the 3,735,289,856-parameter LoRA plus the head, which sum to 29,365,153,792 (reasoned). Merge the LoRA, as you would to serve it, and it is a 25.6B model with a 5-million-parameter head.
What the evaluation leaves out

Each card has three columns: JevArena, a "human-labelled transfer" median over 15 tasks, and the Jev Decision Index. Every size leads the same-size models it chose to list (reported): Kai 48.6 on JevArena against GLiNER2.5-Decide's 42.5, Nox 63.6 against Decider 4B's 61.9, Vega 74.0 against AutoJev-27B's 72.1. The cards are careful about the close calls — Nox and Vega are "statistically level" with the runner-up, their words.
What none of the six tables has is a row for Jev. The nearest you can get is to line up numbers from different runners on the same Index version (all reported): Vega's 56.5, Liquid's run putting Jev at 57.9 and d1 at 58.9, and the decider README placing a stock Gemma-4-31B read through its letter readout at 57.33. On that basis the largest Decision 2.0 model sits just under Jev and just under a stock checkpoint with a temperature fitted per option count (reasoned, from three separate runs). The Pareto figure says the same thing in its own picture.
Smaller things: Lux's card gives 46.3 on the Index while Vega's card puts Lux at 45.3 (measured); "training data audited at row level against all Index test items" is a contamination claim I cannot check without the data; and latencies are medians "on a single GPU," unnamed, from 4.9 ms for Kai to 71.4 ms for Vega (reported).
- architecture
- Decision2Model
- task
- feature-extraction
- library
- transformers
- license
- apache-2.0
- safetensors
- 2 shards
- largest file
- 14.94 GB
- files
- 35
- downloads
- 389
- likes
- 55
A rank-512, alpha-1024 LoRA (3,735,289,856 FP32 parameters, a uniform soup of several adapters) for Qwen3.8-27B, plus a 5,263,872-parameter endpoint head. The card's 29.37B counts the unmerged adapter. Apache-2.0.
repo last modified 2026-10-03
FluidUse: Decision 2.0 on a Mac, at 15 issues a second
FluidInference's port is the week's best engineering, and as with their GLiNER-2.5 port and SigLIP 2 port — both FluidUse, both covered — the write-up is precise about what it checked.

The port. Kai and Eos become fp16 Core ML packages of 1.1 GB and 1.4 GB that run on the GPU; the Neural Engine was slower, 72 ms against 28 ms on Kai's quickstart (reported). The clever part is the packing. All questions of a request go in one call: the shared context is the prefix, each question's tokens follow, and the mask lets a question see the prefix and itself but not the others, so "every question therefore sees exactly the prompt it would see alone." Eos is harder, because its Gated DeltaNet layers carry a recurrent state rather than a cache; the port runs the prefix once and restarts each question from its final state.
Parity, measured properly. On 400 requests (2,000 decisions) from the
typed-decisions test split: 0 token mismatches, 5 flipped answers for Kai and
4 for Eos, all near-ties with an upstream top-two margin of at most 0.009, and a
largest probability difference of 0.0082 (reported, with the script to
reproduce it). On an M5 Pro a five-question request takes 63 ms on Kai and
113 ms on Eos, against 418 and 1,595 ms for the upstream PyTorch runtime on MPS.
The demo. IssueTriageDemo takes 1,000 real issues from
vllm-project/semantic-router, strips their labels, and asks five questions per
issue in one call: type, owning workgroup, priority, needs-info and
good-first-issue. Its README: 66.8 ms per issue, 15.0 issues a second, 5,000
decisions in 67 seconds (reported). Its check is that the labels "match the
Python Core ML reference demo on every issue compared (626 of 626)" — the port
agrees with itself across languages. It never compares the labels with the
real ones it stripped, so the demo says nothing about whether the triage is
right (measured, from the code and README).
Two details in TriageModel.swift say more (measured). The priority question's
score "ranks issues well but its argmax is P0 for most issues," so the demo
cuts it at this repository's 90th and 50th percentiles (1.452 and 1.251): about
10% of issues come out P0 by construction. And the yes/no flags fire at 0.6, a
hand-set cut.
The one accuracy number. The parity tables are the only accuracy figures in the release, and the cards say plainly they "are not the model card's benchmark protocol." On that split, Kai answers 0.423 of choice questions, 0.628 of yes/no and 0.398 of score questions correctly; Eos 0.430, 0.598 and 0.379 (reported). A yes/no accuracy of 0.628 is not far above a coin's half (reasoned), on a held-out set that is not this model's own benchmark. It is also the only number for these two models that someone other than vllm-sr produced.
- license
- apache-2.0
- largest file
- 1.19 GB
- files
- 23
- downloads
- 0
- likes
- 0
Decision-2.0-Kai-0.6B as one fp16 Core ML multifunction package (1.1 GB, four fixed shapes), GPU-only by measurement. Parity with upstream on 2,000 decisions: 0 token mismatches, 5 near-tie flips. Apache-2.0.
repo last modified 2026-10-05
llama.cpp /v1/systemone: one endpoint, six readouts
The any model can be Jev piece found that
llama.cpp had no decision endpoint, and that the nearest primitive returned the
top-N tokens rather than the ones you asked about. That is now out of date.
ggml-org's blog post of 2 October announces /v1/systemone in the server,
"follows the System One format introduced with TypeSafe's Jev," and lists six
supported models (reported): Julia-1 (144M, mmBERT-small), Laya (421M,
ModernBERT-large), Kev-4B, lev (4B), OpenJev (27B, reads images) and Clef (27B),
at a median 3 ms to 43 ms per question on one RTX PRO 6000.

Its design is the opposite of SGLang's (measured, from the server code and
README). SGLang turns any chat model into a letter readout with a flag.
llama.cpp serves only native decision models whose GGUF carries
<arch>.decision.* metadata — "a language model that classifies with prompts
does not get decisions" — and it implements each family's own readout:
type in common.h | how an option is scored | used by |
|---|---|---|
OPENJEV | logit of one label token per option, at the last prompt token | OpenJev |
LEV | the same; noul read from a rating scale; choice run in two orders and averaged | lev |
NIMBLE | the same, with every question listed in the prompt | Bespoke Nimble |
KEV | dot product of the last token's hidden state with each option's end token | Kev-4B |
LAYA | one output column per question type, at each option's marker token | Laya, Julia-1 |
CLEF | all questions in one prompt; option i's score at output row i | Clef |
Decision 2.0's bilinear-plus-MLP head has no type here, so none of this week's open family runs on llama.cpp yet (measured).
After the scores, every family goes through the same format_answer: divide by
a temperature stored in the GGUF (per question type and, where the file
provides them, per option-count bucket), softmax, average the variants, and report. choice
confidence is TypeSafe's formula, ; score is the
expected level. The blog's example response checks against it exactly
(measured): billing at 0.9049 of three options gives confidence
; the urgency probabilities 0.036, 0.1937,
0.2225 and 0.5478 give an expected level of 2.2821 and a score confidence of
0.2821. That last pair is not a coincidence: with the mode on the top of four
levels, the score confidence reduces to the score minus two (reasoned, from the
formula).
as listed
answer
"choice": "billing",
"probabilities": {
"billing": 0.9049,
"shipping": 0.0275,
"technical": 0.0676
},
"confidence": 0.8573"One forward pass" is true with qualifications, and the code is explicit about
them. Questions are answered independently, each in its own prompt, except on
Clef, which decides them jointly; for Kev, lev, OpenJev and Nimble the server
groups the prompts so the shared state is evaluated once. And lev's choice
runs twice, the second time with the options reversed, "to cancel the
preference for the first label." The widget above shows what that does and does
not cancel (reasoned): a first-slot bonus lands on the first option in one order
and the last in the other, so averaging splits it between the two ends rather
than removing it, and with three or more options the middle loses.
The choice cap depends on the model: "52 for openjev, 255 for laya and
clef" (measured, README). And the server's docs end the way SGLang's did: "The probabilities are scaled with the temperatures stored
in the model file. They are not guaranteed to be calibrated for your data." The
blog makes the same point by example: a vague ticket "scored 0.25 with Julia-1
but 0.80 with Kev-4B." A cutoff belongs to a model, not to the endpoint.
What week four measured
Deployment is solved four ways: a closed API, Transformers remote code,
Core ML, and llama.cpp — and d1, Decision 2.0's system_one() and llama.cpp
all answer in the same System One JSON (reasoned).
The receipts came from the closed lab. Liquid published the file behind its chart, which items were written late, and the per-line thresholds its demo needs. Decision 2.0's comparisons stop at same-size open models. FluidUse proved its port faithful, not right. llama.cpp claimed only speed.
Calibration is still the gap. Liquid's docs call d1's outputs calibrated and publish no reliability number, and its own demo fits a threshold per line. Decision 2.0 ships raw temperature-1 probabilities. llama.cpp applies whatever temperature the GGUF carries and disclaims the result. Nobody this week shipped an ECE.
If you are choosing this week
All reasoned, from the evidence above:
- Images, and you do not need the weights: d1, with the per-line threshold work its own demo did, and remembering each question re-bills each image.
- Open weights at a size you can pick: Decision 2.0, Apache-2.0 from 0.6B to 27B, one head design. Measure it against Jev on your data, because its cards do not.
- On a Mac: FluidUse's Kai port, 63 ms for five questions, faithful to upstream. Check accuracy on your labels; the demo did not.
- Local, behind the Jev API, today: llama.cpp with Kev-4B, Laya, lev, OpenJev or Clef. Set the confidence cutoff per model.
The take
Week four turned decision models into infrastructure: an image-reading API, six open sizes, a native Mac runtime and a llama.cpp endpoint, all within days. The best-documented claim of the week came with its own JSON file, and checking it took ten minutes: the averages are exact, the "four of six" is generous by half a point, and the "never trained for this" inspection result needs a threshold fitted per production line.
The open releases did the expensive engineering and skipped the expensive comparison: six sizes with no Jev row, a fast demo with no accuracy, an endpoint that tells you to calibrate. All useful. None, yet, tells you how often it is right.
Sources, read on 2026-10-06: Liquid AI's d1 launch post, decision-model docs and the d1 Playground with its demos/data/compare.json and demos/data/inspection/index.json; the Decision 2.0 collection, including vllm-sr/Decision-2.0-Kai-0.6B and vllm-sr/Decision-2.0-Vega-27B, and vLLM Semantic Router; FluidInference's FluidUse (branch feat/decision-2.0 at e080acd), decision-2.0-kai-coreml and decision-2.0-eos-coreml; ggml-org's decision models in llama.cpp and llama.cpp at 43fe9c6 (tools/server/server-decision.cpp, server-context.cpp, common/common.h); and Mapika/decider with decider-2b-vision. Parameter counts are safetensors header and index reads; no third-party code was executed. Figures are reproduced for commentary from Liquid AI's blog and launch posts, the Decision-2.0-Vega-27B model card (Apache-2.0), FluidInference's demo video and the ggml-org blog; the two interactives are original.