# Jev alternatives, week four: the week decision models learned where to run

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/jev-alternatives-week-four
> date: 2026-10-06
> tags: explainer, llm, ecosystem, benchmarks, calibration, multimodal, on-device, inference

[Week two](/articles/jev-alternatives-week-two) was seven open decision models
and the discovery that most "beats Jev" numbers stood on JevBench's public
items. [Week three](/articles/jev-alternatives-week-three) was the most serious
architectures yet and the thinnest evaluations: not one new model on a sealed
or independent board.

Week four is four posts, and none is mainly about a better model. All four
are about **where a decision model runs**: behind a closed API that now accepts images (Liquid's
d1), as open weights in six sizes loaded through Transformers (vLLM Semantic
Router's Decision 2.0), natively on a Mac through Core ML (FluidUse), and
inside llama.cpp's server behind the same `/v1/systemone` endpoint as Jev
itself.

The evaluation story inverts. The one publisher that shipped its raw results
is the closed one. The open family leaves Jev out of its tables; the on-device
port checks that it matches upstream, never whether it is right; llama.cpp
makes no accuracy claim and tells you to calibrate your own cutoffs.

**Measured** means I read it from a file, a safetensors header or source code,
or recomputed it from a published file; **reported** means the publisher's
figure, which I could not re-run; **reasoned** means my inference from the other
two. Code was read, never executed. All checked on 2026-10-06.

## The one idea, again

A decision model answers a schema of typed questions — `choice`, `score`,
`noul` (yes/no) — by returning a probability for every option, read in one
forward pass, with no tokens generated and nothing to parse. The full
treatment is in [what Jev actually is](/articles/jev-system-one-models), [why
any model can be Jev](/articles/any-model-can-be-jev) and [what decision models
cannot do](/articles/what-decision-models-cannot-do); and [Jev is not
deterministic](/articles/jev-is-not-deterministic) is why "one pass" does not
mean "same answer every time."

The open ones still sort into three readout families: a **letter readout**
(options become letters, read the letters' next-token logits), a **pointer**
(each option's end-token hidden state is scored against the question's, as in
[Kev](/articles/kev)), and a **joint scorer** (a trained head scores everything
at once, as in clef). This week (measured, from configs and code): d1 is
closed and undisclosed; Decision 2.0 is a pointer with one head design across
six sizes; FluidUse re-implements that pointer in Core ML; and llama.cpp is not
a model but six readouts behind one endpoint.

## d1 with vision (Liquid): the best receipt of the week is a JSON file

Liquid launched d1 as a text model on 29 September. The week-four post adds
images: "images, text or both as inputs," tested "against GPT-6.1 Sol and
Claude Opus 5.5 on six real applications," matching or beating GPT-6.1 Sol "on
four of them," at "19x to 200x less" cost (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-four/fig1.png"
  alt="Liquid's diagram of how d1 works. An INPUT panel holds three boxes: a camera photo of a small blue circuit board captioned 'Camera image of a circuit board', 'Question: What is the defect?', and 'Options: missing, extra, scratch'. Lines from all three join into a black box labelled d1, which points to an OUTPUT panel titled Probabilities: missing 0.80, extra 0.11, scratch 0.10."
  caption="The interface, not the mechanism: an image, a question and its options go in, a probability per option comes out. Liquid publishes nothing about what sits inside the black box. (Liquid AI, 'Introducing d1', Figure 2.)"
/>

What can be checked is the interface. The request goes to
`https://api.liquid.ai/decisions/v1/systemone`, the docs install TypeSafe's own
`typesafe-sdk` and point its `base_url` at Liquid, and the question types are
Jev's `noul`, `choice` and `score` — so d1 is a drop-in for Jev's wire format,
the same thing [week three](/articles/jev-alternatives-week-three) watched
SGLang and Unsloth build. Billing is input tokens only, at \$0.04 per million.
An image costs 1.5 tokens per 32×32 patch, so a 1024×1024 image is 1,536
tokens, and "every question is charged for its text and all the images again"
(measured, from the docs). That last clause matters for vision: ask five
questions about one photo and you pay for the photo five times. Requests take
up to 8 images and 10,000 patches.

### The chart, and the file behind it

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-four/fig2.jpg"
  alt="Liquid's launch chart titled 'd1 against frontier models, averaged over six applications', three bar groups. Quality: d1 94%, GPT-6.1 Sol 91%, Claude Opus 5.5 98%. Cost per 1,000 runs: d1 $0.54, GPT-6.1 Sol $26, Claude Opus 5.5 $85. Time per run: d1 3.7 s, GPT-6.1 Sol 14.1 s, Claude Opus 5.5 15.8 s."
  caption="The launch chart. Every bar is an average over six applications whose 'runs' are different units: a coding session, a passage, a 150-row query, a photo, a question, a goal. (Liquid AI, d1 launch post on X.)"
/>

The d1 playground is a static site, and its comparison panel loads
`demos/data/compare.json`: per application and per model, quality, cost,
seconds, request counts and token counts, stamped "measured 2026-10-05". I
recomputed the chart from it (measured):

- **Quality** averages to 93.6%, 91.3% and 98.2% — the chart's 94, 91 and 98.
- **Time** averages to 3.74, 14.15 and 15.75 seconds — the chart's 3.7, 14.1
  and 15.8.
- **Cost** reproduces only one way. With Smart Folders' cost taken per passage
  (the file stores it per 105 passages) and every other application per run,
  the means are \$0.54, \$25.5 and \$84.7 per 1,000 runs — the chart's \$0.54,
  \$26 and \$85. The blog's methodology confirms that Smart Folders' "cost is
  per 1,000 passages."

So the chart is honest arithmetic over an odd unit: "cost per 1,000 runs"
averages a coding session, a passage, a 150-ticket query, a photo, a question
and a web goal, and Code Search and Smart Filter make up most of the frontier
models' averages (reasoned). The per-application cost ratios are
the cleaner number: GPT-6.1 Sol costs **18.7x** d1's on context compaction and
up to 56.5x on Smart Folders; Opus costs 61.4x on compaction and **200.1x** on
Code Search (measured, from the file). That is exactly the post's "19x to
200x."

<D1Ledger />

### "Matches or beats GPT-6.1 Sol on four"

Per application (measured, from the file):

| application | d1 | GPT-6.1 Sol | Claude Opus 5.5 | sample |
|---|---:|---:|---:|---|
| Context Compaction | 7 of 7 | 6 of 7 | 7 of 7 | 7 needed outputs |
| Visual Inspection | 90.6% | 82.3% | 91.7% | 96 photos |
| Web Agent | 7 of 7 | 7 of 7 | 7 of 7 | 7 goals |
| Smart Filter (F1) | 94.7 | 95.2 | 97.6 | 150 tickets |
| Smart Folders | 96.2% | 98.1% | 100% | 105 passages |
| Code Search | 12 of 15 | 13 of 15 | 15 of 15 | 15 questions |

d1 beats GPT-6.1 Sol on two (compaction, inspection), ties on one (web), and
loses on three. The fourth "match" is Smart Filter, where the file's own
display string rounds both F1 scores to "95%" — 94.7 against 95.2 unrounded.
That is within any reasonable noise for 150 tickets, so "matches" is
defensible; it is also the generous reading. Against Opus, d1 ties on
compaction and web and trails on the other four.

Each application was run once, on seven to 150 items. Liquid's methodology
note adds something most launch posts never do: "We wrote six of the 15 code
questions and two of the four compaction sessions after d1's pipeline was set."
That is a disclosure in d1's favour — those items were not tuned against — and
a reminder of how small the sets are. The chat models answered in batches (48
requests for Smart Filter against d1's 900) at list prices without prompt
caching. The chart's 3.7 seconds is per run with "up to 8 requests in flight";
the post's "text decisions in 200 to 300 ms" is a separate claim this file
cannot check (reported).

### The inspection result, and the threshold it needs

"85% to 97% accuracy across four VisA inspection tasks… without task-specific
training" (reported). The weights were not trained on VisA. The decision cut
was fitted per line. The playground's `inspection/index.json` stores, for each
production line, a threshold "calibrated at … on … good and … defective parts
that are not shown here" (measured, from the file and the page code):

| line | threshold on `noul` | calibration parts | balanced accuracy at calibration |
|---|---:|---|---:|
| circuit boards | 0.484 | 98 good, 88 defective | 0.952 |
| candles | 0.453 | 98 good, 87 defective | 0.902 |
| cashews | 0.126 | 48 good, 88 defective | 0.900 |
| chewing gum | 0.725 | 49 good, 88 defective | 0.960 |

Two things follow (reasoned). First, a `noul` of 0.5 does not mean the same
thing on every line: the cashew line calls a part defective at 0.126, the gum
line only at 0.725. Liquid's docs say the model "returns calibrated
probabilities"; a per-line cut ranging from 0.126 to 0.725 says that, for
defect detection at least, you still fit your own threshold on labelled
examples, and the demo did. Second, every model in the comparison also sees "a
good part from the same line next to the part to inspect," so this is
reference-based inspection, not zero-shot anomaly detection. The 85–97% range
itself I could not reproduce from these files: the calibration accuracies run
90–96%, and the comparison file's d1 figure over all 96 photos is 90.6%.

The games are reported and uncheckable without the harness: Tetris from 70 to
81 cleared lines when the screen is added to a text description, Wordle 12 of 12
solved in 3.8 guesses from screenshots, and Quick, Draw! at 5.2 of 6 doodles
named among 62 words (random: 0.6).

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-four/fig6.jpg"
  alt="Liquid's chart 'Liquid AI's first decision model: d1', comparing Liquid d1 and Jev 1.13 by area on the Decision Index: Arts 45.5 vs 37.7 (+7.8), Language 67.6 vs 62.0 (+5.6), Retrieval 60.7 vs 55.4 (+5.3), Tools 74.1 vs 75.1 (−1.0), Knowledge 43.3 vs 51.3 (−8.0), Index over 5 areas 58.9 vs 57.9 (+1.0). Footer: Internal reproduction of Decision Index 0.2.1 by Hugging Face."
  caption="d1's only head-to-head with Jev, from the text launch a week earlier: 58.9 against 57.9 on the Decision Index, with d1 behind on Tools and eight points behind on Knowledge. 'Internal reproduction' means Liquid ran it. (Liquid AI, d1 text launch post on X.)"
/>

The blog calls d1 "the first to extend these capabilities to images." It is
not: [Jev-Omni](/articles/jev-omni) went
multimodal first, clef and JEMM read images in week three, and a reply to the
launch post points at `Mapika/decider-2b-vision`, an Apache-2.0 vision decision
model on Hugging Face since 16 September (measured, from the Hub API). Its card
reports Visual7W 0.89. The one comparison to Jev, the chart above, is Liquid's own
run, and the release has no calibration number. Open weights are promised only
for "upcoming models."

## Decision 2.0 (vLLM Semantic Router): six sizes, one head

The post is one line: "our newest state-of-the-art decision models, open in
every size from 0.6B to 27B." The collection has six, each named, each
Apache-2.0:

| model | base | backbone parameters (measured) | head parameters (measured) |
|---|---|---:|---:|
| Kai-0.6B | Qwen3-0.6B-Base | 596,049,920 | 1,053,184 |
| Eos-0.8B | Decision 1.0 Eos | 752,393,024 | 1,053,184 |
| Sol-2B | Decision 1.0 Sol | 1,881,825,088 | 2,105,856 |
| Nox-4B | Qwen3.5-4B-Base | 4,205,751,296 | 2,632,192 |
| Lux-9B | Decision 1.0 Lux | 7,936,684,544 | 4,211,200 |
| Vega-27B | Qwen3.8-27B + LoRA | rank-512 LoRA, 3,735,289,856 | 5,263,872 |

Backbone counts are safetensors index and header sums; head counts are the
`decision_head.safetensors` headers, read by range request.

<ModelCard repo="vllm-sr/Decision-2.0-Kai-0.6B" note="Every weight of Qwen3-0.6B-Base fine-tuned plus a new 1,053,184-parameter FP32 endpoint head; the release is a uniform average of six fine-tunes (two recipes, three seeds). Raw probabilities at temperature 1, no calibration file. Apache-2.0." />

### The head, from the code

The repositories ship their runtime as remote code, and
`decision2/_vendor/dev2model/decision_model.py` is short enough to read whole
(measured). Each question becomes its own prompt:

```text
Context:
<state>

Task type: choice
Question:
<instructions>
Options:
<option>
{"key": "returns", "description": "Refunds, replacements and damaged deliveries"}
</option>
...

Select the single option best supported by the context and instructions.
Decision:
```

The model runs that through the backbone once. It keeps the last hidden state
of each `</option>` block (the "candidate endpoints") and the hidden state of
the final token (the "global query"). The `CandidateHead` scores each option as
a bilinear term plus a small MLP term:

$$
s_i = \frac{(W_k\,\tilde c_i)\cdot(W_q\,\tilde q)}{\sqrt{256}} + w^\top \mathrm{GELU}(W_c\,\tilde c_i + W_{q'}\,\tilde q)
$$

where $\tilde c_i$ and $\tilde q$ are the layer-normed option and query states
and every projection is 256 wide. A softmax over the $s_i$ is the answer. The
head is identical across sizes; only its input width changes (1,024 for Kai up
to 5,120 for Vega), which is why it grows from about one to about five million
parameters.

This is the pointer family. Because the backbone is causal, an option sees the
options listed before it and not those after, so option order can still move
the answer; there is no reverse-and-average in the runtime. The cap is 255
options. Every question is a separate prompt, but the questions share their
context prefix, which is what lets the Core ML port below run them together.

Two more things the files say (measured). `config.json` carries `"calibration":
null`, and the manifest calls the output "raw probabilities (temperature 1; no
calibration file)," with fixed offsets for five-level `score` questions; the
loss has an optional Brier term, so calibration is something training may push
toward, not something fitted after. And the releases are soups: Kai averages six
full fine-tunes, and Vega averages LoRA adapters with their "factors
concatenated along the rank," which is how a rank-512 adapter holds 3.7 billion
parameters.

The 27B's card says "29.37B" parameters. That is the Qwen3.8-27B text model's
25,624,600,064 plus the 3,735,289,856-parameter LoRA plus the head, which sum to
29,365,153,792 (reasoned). Merge the LoRA, as you would to serve it, and it is a
25.6B model with a 5-million-parameter head.

### What the evaluation leaves out

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-four/fig3.png"
  alt="Line chart titled 'Jev Decision Index: balanced skill against model size, Decision 2.0 vs. Decision 1.0 and 64 public entrants', log-scale parameters on x from 0.1B to 30B, balanced skill 0 to 70 on y. A blue Decision 2.0 line rises from about 16 at 0.6B to 56.5 at Decision-2.0-Vega-27B, circled; Decision-2.0-Lux-9B is labelled 45.3. A grey Decision 1.0 line sits below. Grey dots are public entrants; a dashed Pareto frontier runs at about 57 from just under 27B to the right edge, above the Vega point. Footer: Decision 2.0 independent reproduction with the official 0.2.1 kit on the released weights; others public board snapshot 2026-09-28. Training data audited at row level against all Index test items."
  caption="Decision 2.0's own placement on the Decision Index. Vega's 56.5 is circled, and the dashed Pareto frontier passes above it: at least one public entrant at a similar size scores higher. This figure labels Lux 45.3; Lux's own card says 46.3. (vllm-sr/Decision-2.0-Vega-27B model card.)"
/>

Each card has three columns: JevArena, a "human-labelled transfer" median over
15 tasks, and the Jev Decision Index. Every size leads the same-size models it
chose to list (reported): Kai 48.6 on JevArena against GLiNER2.5-Decide's 42.5,
Nox 63.6 against Decider 4B's 61.9, Vega 74.0 against AutoJev-27B's 72.1. The
cards are careful about the close calls — Nox and Vega are "statistically
level" with the runner-up, their words.

What none of the six tables has is a row for Jev. The nearest you can get is to
line up numbers from different runners on the same Index version (all
reported): Vega's 56.5, Liquid's run putting Jev at 57.9 and d1 at 58.9, and
the decider README placing a stock Gemma-4-31B read through its letter readout
at 57.33. On that basis the largest Decision 2.0 model sits just under Jev and
just under a stock checkpoint with a temperature fitted per option count
(reasoned, from three separate runs). The Pareto figure says the same thing in
its own picture.

Smaller things: Lux's card gives 46.3 on the Index while Vega's card puts Lux
at 45.3 (measured); "training data audited at row level against all Index test
items" is a contamination claim I cannot check without the data; and latencies
are medians "on a single GPU," unnamed, from 4.9 ms for Kai to 71.4 ms for Vega
(reported).

<ModelCard repo="vllm-sr/Decision-2.0-Vega-27B" note="A rank-512, alpha-1024 LoRA (3,735,289,856 FP32 parameters, a uniform soup of several adapters) for Qwen3.8-27B, plus a 5,263,872-parameter endpoint head. The card's 29.37B counts the unmerged adapter. Apache-2.0." />

## FluidUse: Decision 2.0 on a Mac, at 15 issues a second

FluidInference's port is the week's best engineering, and as with their
[GLiNER-2.5 port](/articles/gliner-2-5-decide) and [SigLIP 2
port](/articles/siglip-2-coreml) — both FluidUse, both covered — the write-up is
precise about what it checked.

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-four/fig4.jpg"
  alt="A frame from FluidInference's demo video. Left: a dark GitHub-style issues page for 'routerlabs / semantic-switch' with a purple 'Triage with Decision 2.0' button, a stats bar reading '787 / 1000 issues triaged, 3935 labels decided, 72 ms per issue, 13.3 issues / sec' and '5 decisions per issue in one call · Decision-2.0-Kai-0.6B · Core ML on this Mac', a sidebar of workgroup counts, and issues gaining labels such as enhancement, wg/mom-routing and priority/P2. Right: a terminal showing Apple M5 Pro usage with GPU at 97% and ANE at 0%, and per-issue log lines of about 71 ms each."
  caption="Mid-run: 787 of 1,000 issues, five labels each, at 72 ms per issue. The system monitor on the right shows where the work runs — GPU at 97%, the Neural Engine at 0%. (FluidInference, FluidUse demo video on X.)"
/>

**The port.** Kai and Eos become fp16 Core ML packages of 1.1 GB and 1.4 GB
that run on the GPU; the Neural Engine was slower, 72 ms against 28 ms on Kai's
quickstart (reported). The clever part is the packing. All questions of a
request go in one call: the shared context is the prefix, each question's
tokens follow, and the mask lets a question see the prefix and itself but not
the others, so "every question therefore sees exactly the prompt it would see
alone." Eos is harder, because its Gated DeltaNet layers carry a recurrent state
rather than a cache; the port runs the prefix once and restarts each question
from its final state.

**Parity, measured properly.** On 400 requests (2,000 decisions) from the
`typed-decisions` test split: 0 token mismatches, 5 flipped answers for Kai and
4 for Eos, all near-ties with an upstream top-two margin of at most 0.009, and a
largest probability difference of 0.0082 (reported, with the script to
reproduce it). On an M5 Pro a five-question request takes 63 ms on Kai and
113 ms on Eos, against 418 and 1,595 ms for the upstream PyTorch runtime on MPS.

**The demo.** `IssueTriageDemo` takes 1,000 real issues from
`vllm-project/semantic-router`, strips their labels, and asks five questions per
issue in one call: type, owning workgroup, priority, needs-info and
good-first-issue. Its README: 66.8 ms per issue, 15.0 issues a second, 5,000
decisions in 67 seconds (reported). Its check is that the labels "match the
Python Core ML reference demo on every issue compared (626 of 626)" — the port
agrees with itself across languages. It never compares the labels with the
real ones it stripped, so the demo says nothing about whether the triage is
right (measured, from the code and README).

Two details in `TriageModel.swift` say more (measured). The priority question's
`score` "ranks issues well but its argmax is P0 for most issues," so the demo
cuts it at this repository's 90th and 50th percentiles (1.452 and 1.251): about
10% of issues come out P0 by construction. And the yes/no flags fire at 0.6, a
hand-set cut.

**The one accuracy number.** The parity tables are the only accuracy figures in
the release, and the cards say plainly they "are not the model card's
benchmark protocol." On that split, Kai answers 0.423 of choice questions, 0.628
of yes/no and 0.398 of score questions correctly; Eos 0.430, 0.598 and 0.379
(reported). A yes/no accuracy of 0.628 is not far above a coin's half (reasoned),
on a held-out set that is not this model's own benchmark. It is also the only
number for these two models that someone other than vllm-sr produced.

<ModelCard repo="FluidInference/decision-2.0-kai-coreml" note="Decision-2.0-Kai-0.6B as one fp16 Core ML multifunction package (1.1 GB, four fixed shapes), GPU-only by measurement. Parity with upstream on 2,000 decisions: 0 token mismatches, 5 near-tie flips. Apache-2.0." />

## llama.cpp `/v1/systemone`: one endpoint, six readouts

The [any model can be Jev](/articles/any-model-can-be-jev) piece found that
llama.cpp had no decision endpoint, and that the nearest primitive returned the
top-N tokens rather than the ones you asked about. That is now out of date.
ggml-org's blog post of 2 October announces `/v1/systemone` in the server,
"follows the System One format introduced with TypeSafe's Jev," and lists six
supported models (reported): Julia-1 (144M, mmBERT-small), Laya (421M,
ModernBERT-large), Kev-4B, lev (4B), OpenJev (27B, reads images) and Clef (27B),
at a median 3 ms to 43 ms per question on one RTX PRO 6000.

<Figure
  src="https://ai.thesatyajit.com/articles/jev-alternatives-week-four/fig5.png"
  alt="The header graphic of the ggml-org blog post: the llama.cpp wordmark on a dark background, beside a tree of nodes in which one path from the root is drawn in orange to a single highlighted leaf among five."
  caption="The post's header: one path chosen among several. The post has no architecture diagram; the readout lives in tools/server/server-decision.cpp. (ggml-org, 'New in llama.cpp: Decision Models'.)"
/>

Its design is the opposite of SGLang's (measured, from the server code and
README). SGLang turns any chat model into a letter readout with a flag.
llama.cpp serves only **native** decision models whose GGUF carries
`<arch>.decision.*` metadata — "a language model that classifies with prompts
does not get `decisions`" — and it implements each family's own readout:

| type in `common.h` | how an option is scored | used by |
|---|---|---|
| `OPENJEV` | logit of one label token per option, at the last prompt token | OpenJev |
| `LEV` | the same; `noul` read from a rating scale; `choice` run in two orders and averaged | lev |
| `NIMBLE` | the same, with every question listed in the prompt | Bespoke Nimble |
| `KEV` | dot product of the last token's hidden state with each option's end token | Kev-4B |
| `LAYA` | one output column per question type, at each option's marker token | Laya, Julia-1 |
| `CLEF` | all questions in one prompt; option *i*'s score at output row *i* | Clef |

Decision 2.0's bilinear-plus-MLP head has no type here, so none of this week's
open family runs on llama.cpp yet (measured).

After the scores, every family goes through the same `format_answer`: divide by
a temperature stored in the GGUF (per question type and, where the file
provides them, per option-count bucket), softmax, average the variants, and report. `choice`
confidence is TypeSafe's formula, $(p_{max} - 1/n)/(1 - 1/n)$; `score` is the
expected level. The blog's example response checks against it exactly
(measured): billing at 0.9049 of three options gives confidence
$(0.9049 - 1/3)/(2/3) = 0.8574$; the urgency probabilities 0.036, 0.1937,
0.2225 and 0.5478 give an expected level of 2.2821 and a score confidence of
0.2821. That last pair is not a coincidence: with the mode on the top of four
levels, the score confidence reduces to the score minus two (reasoned, from the
formula).

<ReadoutExplorer />

"One forward pass" is true with qualifications, and the code is explicit about
them. Questions are answered independently, each in its own prompt, except on
Clef, which decides them jointly; for Kev, lev, OpenJev and Nimble the server
groups the prompts so the shared state is evaluated once. And lev's `choice`
runs **twice**, the second time with the options reversed, "to cancel the
preference for the first label." The widget above shows what that does and does
not cancel (reasoned): a first-slot bonus lands on the first option in one order
and the last in the other, so averaging splits it between the two ends rather
than removing it, and with three or more options the middle loses.

The `choice` cap depends on the model: "52 for openjev, 255 for laya and
clef" (measured, README). And the server's docs end the way SGLang's did: "The probabilities are scaled with the temperatures stored
in the model file. They are not guaranteed to be calibrated for your data." The
blog makes the same point by example: a vague ticket "scored 0.25 with Julia-1
but 0.80 with Kev-4B." A cutoff belongs to a model, not to the endpoint.

## What week four measured

**Deployment is solved four ways**: a closed API, Transformers remote code,
Core ML, and llama.cpp — and d1, Decision 2.0's `system_one()` and llama.cpp
all answer in the same System One JSON (reasoned).

**The receipts came from the closed lab.** Liquid published the file behind
its chart, which items were written late, and the per-line thresholds its demo
needs. Decision 2.0's comparisons stop at same-size open models. FluidUse proved
its port faithful, not right. llama.cpp claimed only speed.

**Calibration is still the gap.** Liquid's docs call d1's outputs calibrated
and publish no reliability number, and its own demo fits a threshold per line.
Decision 2.0 ships raw temperature-1 probabilities. llama.cpp applies whatever
temperature the GGUF carries and disclaims the result. Nobody this week shipped
an ECE.

## If you are choosing this week

All reasoned, from the evidence above:

- **Images, and you do not need the weights:** d1, with the per-line threshold
  work its own demo did, and remembering each question re-bills each image.
- **Open weights at a size you can pick:** Decision 2.0, Apache-2.0 from 0.6B to
  27B, one head design. Measure it against Jev on your data, because its cards
  do not.
- **On a Mac:** FluidUse's Kai port, 63 ms for five questions, faithful to
  upstream. Check accuracy on your labels; the demo did not.
- **Local, behind the Jev API, today:** llama.cpp with Kev-4B, Laya, lev,
  OpenJev or Clef. Set the confidence cutoff per model.

## The take

Week four turned decision models into infrastructure: an image-reading API, six
open sizes, a native Mac runtime and a llama.cpp endpoint, all within days.
The best-documented claim of the week came with its own JSON file, and checking
it took ten minutes: the averages are exact, the "four of six" is generous by
half a point, and the "never trained for this" inspection result needs a
threshold fitted per production line.

The open releases did the expensive engineering and skipped the expensive
comparison: six sizes with no Jev row, a fast demo with no accuracy, an
endpoint that tells you to calibrate. All useful. None, yet, tells you how often
it is right.

---

*Sources, read on 2026-10-06: Liquid AI's [d1 launch post](https://www.liquid.ai/blog/d1-decision-model), [decision-model docs](https://docs.liquid.ai/lfm/models/decision-models) and the [d1 Playground](https://d1.liquid.ai/) with its `demos/data/compare.json` and `demos/data/inspection/index.json`; the [Decision 2.0 collection](https://huggingface.co/collections/vllm-sr/decision-20), including [vllm-sr/Decision-2.0-Kai-0.6B](https://huggingface.co/vllm-sr/Decision-2.0-Kai-0.6B) and [vllm-sr/Decision-2.0-Vega-27B](https://huggingface.co/vllm-sr/Decision-2.0-Vega-27B), and [vLLM Semantic Router](https://github.com/vllm-project/semantic-router); FluidInference's [FluidUse](https://github.com/FluidInference/FluidUse) (branch `feat/decision-2.0` at `e080acd`), [decision-2.0-kai-coreml](https://huggingface.co/FluidInference/decision-2.0-kai-coreml) and [decision-2.0-eos-coreml](https://huggingface.co/FluidInference/decision-2.0-eos-coreml); ggml-org's [decision models in llama.cpp](https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp) and [llama.cpp](https://github.com/ggml-org/llama.cpp) at `43fe9c6` (`tools/server/server-decision.cpp`, `server-context.cpp`, `common/common.h`); and [Mapika/decider](https://github.com/Mapika/decider) with [decider-2b-vision](https://huggingface.co/Mapika/decider-2b-vision). Parameter counts are safetensors header and index reads; no third-party code was executed. Figures are reproduced for commentary from Liquid AI's blog and launch posts, the Decision-2.0-Vega-27B model card (Apache-2.0), FluidInference's demo video and the ggml-org blog; the two interactives are original.*
