~/satyajit

Jev alternatives, week four: the week decision models learned where to run

mdjsonmcp

2026-10-06 · 23 min · explainer · llm · ecosystem · benchmarks · calibration · multimodal · on-device · inference

Week two was seven open decision models and the discovery that most "beats Jev" numbers stood on JevBench's public items. Week three was the most serious architectures yet and the thinnest evaluations: not one new model on a sealed or independent board.

Week four is four posts, and none is mainly about a better model. All four are about where a decision model runs: behind a closed API that now accepts images (Liquid's d1), as open weights in six sizes loaded through Transformers (vLLM Semantic Router's Decision 2.0), natively on a Mac through Core ML (FluidUse), and inside llama.cpp's server behind the same /v1/systemone endpoint as Jev itself.

The evaluation story inverts. The one publisher that shipped its raw results is the closed one. The open family leaves Jev out of its tables; the on-device port checks that it matches upstream, never whether it is right; llama.cpp makes no accuracy claim and tells you to calibrate your own cutoffs.

Measured means I read it from a file, a safetensors header or source code, or recomputed it from a published file; reported means the publisher's figure, which I could not re-run; reasoned means my inference from the other two. Code was read, never executed. All checked on 2026-10-06.

The one idea, again

A decision model answers a schema of typed questions — choice, score, noul (yes/no) — by returning a probability for every option, read in one forward pass, with no tokens generated and nothing to parse. The full treatment is in what Jev actually is, why any model can be Jev and what decision models cannot do; and Jev is not deterministic is why "one pass" does not mean "same answer every time."

The open ones still sort into three readout families: a letter readout (options become letters, read the letters' next-token logits), a pointer (each option's end-token hidden state is scored against the question's, as in Kev), and a joint scorer (a trained head scores everything at once, as in clef). This week (measured, from configs and code): d1 is closed and undisclosed; Decision 2.0 is a pointer with one head design across six sizes; FluidUse re-implements that pointer in Core ML; and llama.cpp is not a model but six readouts behind one endpoint.

d1 with vision (Liquid): the best receipt of the week is a JSON file

Liquid launched d1 as a text model on 29 September. The week-four post adds images: "images, text or both as inputs," tested "against GPT-6.1 Sol and Claude Opus 5.5 on six real applications," matching or beating GPT-6.1 Sol "on four of them," at "19x to 200x less" cost (reported).

Liquid's diagram of how d1 works. An INPUT panel holds three boxes: a camera photo of a small blue circuit board captioned 'Camera image of a circuit board', 'Question: What is the defect?', and 'Options: missing, extra, scratch'. Lines from all three join into a black box labelled d1, which points to an OUTPUT panel titled Probabilities: missing 0.80, extra 0.11, scratch 0.10.
The interface, not the mechanism: an image, a question and its options go in, a probability per option comes out. Liquid publishes nothing about what sits inside the black box. (Liquid AI, 'Introducing d1', Figure 2.)

What can be checked is the interface. The request goes to https://api.liquid.ai/decisions/v1/systemone, the docs install TypeSafe's own typesafe-sdk and point its base_url at Liquid, and the question types are Jev's noul, choice and score — so d1 is a drop-in for Jev's wire format, the same thing week three watched SGLang and Unsloth build. Billing is input tokens only, at $0.04 per million. An image costs 1.5 tokens per 32×32 patch, so a 1024×1024 image is 1,536 tokens, and "every question is charged for its text and all the images again" (measured, from the docs). That last clause matters for vision: ask five questions about one photo and you pay for the photo five times. Requests take up to 8 images and 10,000 patches.

The chart, and the file behind it

Liquid's launch chart titled 'd1 against frontier models, averaged over six applications', three bar groups. Quality: d1 94%, GPT-6.1 Sol 91%, Claude Opus 5.5 98%. Cost per 1,000 runs: d1 $0.54, GPT-6.1 Sol $26, Claude Opus 5.5 $85. Time per run: d1 3.7 s, GPT-6.1 Sol 14.1 s, Claude Opus 5.5 15.8 s.
The launch chart. Every bar is an average over six applications whose 'runs' are different units: a coding session, a passage, a 150-row query, a photo, a question, a goal. (Liquid AI, d1 launch post on X.)

The d1 playground is a static site, and its comparison panel loads demos/data/compare.json: per application and per model, quality, cost, seconds, request counts and token counts, stamped "measured 2026-10-05". I recomputed the chart from it (measured):

So the chart is honest arithmetic over an odd unit: "cost per 1,000 runs" averages a coding session, a passage, a 150-ticket query, a photo, a question and a web goal, and Code Search and Smart Filter make up most of the frontier models' averages (reasoned). The per-application cost ratios are the cleaner number: GPT-6.1 Sol costs 18.7x d1's on context compaction and up to 56.5x on Smart Folders; Opus costs 61.4x on compaction and 200.1x on Code Search (measured, from the file). That is exactly the post's "19x to 200x."

d1 vs two frontier chat models · Liquid's own run, one application at a time
one run
cost: a mean over six different kinds of run; time: six different kinds of run
quality
mean of six different metrics
sample
six applications

quality (axis from 75%)

d1
94%
GPT-6.1 Sol
91%
Claude Opus 5.5
98%

cost per 1,000 runs (log axis)

d1
$0.54
GPT-6.1 Sol
$25.5
Claude Opus 5.5
$84.7

seconds per run

d1
3.7 s
GPT-6.1 Sol
14.1 s
Claude Opus 5.5
15.8 s

d1 is 47.2x cheaper than GPT-6.1 Sol and 156.5x cheaper than Claude Opus 5.5 here. Across single applications the range is 18.7x to 200.1x.

Values from Liquid's published comparison file, run once per model on 2026-10-05; nothing here was re-run. The average row reproduces Liquid's chart, and its cost column is a mean over six differently sized runs, dominated by Code Search and Smart Filter.

"Matches or beats GPT-6.1 Sol on four"

Per application (measured, from the file):

applicationd1GPT-6.1 SolClaude Opus 5.5sample
Context Compaction7 of 76 of 77 of 77 needed outputs
Visual Inspection90.6%82.3%91.7%96 photos
Web Agent7 of 77 of 77 of 77 goals
Smart Filter (F1)94.795.297.6150 tickets
Smart Folders96.2%98.1%100%105 passages
Code Search12 of 1513 of 1515 of 1515 questions

d1 beats GPT-6.1 Sol on two (compaction, inspection), ties on one (web), and loses on three. The fourth "match" is Smart Filter, where the file's own display string rounds both F1 scores to "95%" — 94.7 against 95.2 unrounded. That is within any reasonable noise for 150 tickets, so "matches" is defensible; it is also the generous reading. Against Opus, d1 ties on compaction and web and trails on the other four.

Each application was run once, on seven to 150 items. Liquid's methodology note adds something most launch posts never do: "We wrote six of the 15 code questions and two of the four compaction sessions after d1's pipeline was set." That is a disclosure in d1's favour — those items were not tuned against — and a reminder of how small the sets are. The chat models answered in batches (48 requests for Smart Filter against d1's 900) at list prices without prompt caching. The chart's 3.7 seconds is per run with "up to 8 requests in flight"; the post's "text decisions in 200 to 300 ms" is a separate claim this file cannot check (reported).

The inspection result, and the threshold it needs

"85% to 97% accuracy across four VisA inspection tasks… without task-specific training" (reported). The weights were not trained on VisA. The decision cut was fitted per line. The playground's inspection/index.json stores, for each production line, a threshold "calibrated at … on … good and … defective parts that are not shown here" (measured, from the file and the page code):

linethreshold on noulcalibration partsbalanced accuracy at calibration
circuit boards0.48498 good, 88 defective0.952
candles0.45398 good, 87 defective0.902
cashews0.12648 good, 88 defective0.900
chewing gum0.72549 good, 88 defective0.960

Two things follow (reasoned). First, a noul of 0.5 does not mean the same thing on every line: the cashew line calls a part defective at 0.126, the gum line only at 0.725. Liquid's docs say the model "returns calibrated probabilities"; a per-line cut ranging from 0.126 to 0.725 says that, for defect detection at least, you still fit your own threshold on labelled examples, and the demo did. Second, every model in the comparison also sees "a good part from the same line next to the part to inspect," so this is reference-based inspection, not zero-shot anomaly detection. The 85–97% range itself I could not reproduce from these files: the calibration accuracies run 90–96%, and the comparison file's d1 figure over all 96 photos is 90.6%.

The games are reported and uncheckable without the harness: Tetris from 70 to 81 cleared lines when the screen is added to a text description, Wordle 12 of 12 solved in 3.8 guesses from screenshots, and Quick, Draw! at 5.2 of 6 doodles named among 62 words (random: 0.6).

Liquid's chart 'Liquid AI's first decision model: d1', comparing Liquid d1 and Jev 1.13 by area on the Decision Index: Arts 45.5 vs 37.7 (+7.8), Language 67.6 vs 62.0 (+5.6), Retrieval 60.7 vs 55.4 (+5.3), Tools 74.1 vs 75.1 (−1.0), Knowledge 43.3 vs 51.3 (−8.0), Index over 5 areas 58.9 vs 57.9 (+1.0). Footer: Internal reproduction of Decision Index 0.2.1 by Hugging Face.
d1's only head-to-head with Jev, from the text launch a week earlier: 58.9 against 57.9 on the Decision Index, with d1 behind on Tools and eight points behind on Knowledge. 'Internal reproduction' means Liquid ran it. (Liquid AI, d1 text launch post on X.)

The blog calls d1 "the first to extend these capabilities to images." It is not: Jev-Omni went multimodal first, clef and JEMM read images in week three, and a reply to the launch post points at Mapika/decider-2b-vision, an Apache-2.0 vision decision model on Hugging Face since 16 September (measured, from the Hub API). Its card reports Visual7W 0.89. The one comparison to Jev, the chart above, is Liquid's own run, and the release has no calibration number. Open weights are promised only for "upcoming models."

Decision 2.0 (vLLM Semantic Router): six sizes, one head

The post is one line: "our newest state-of-the-art decision models, open in every size from 0.6B to 27B." The collection has six, each named, each Apache-2.0:

modelbasebackbone parameters (measured)head parameters (measured)
Kai-0.6BQwen3-0.6B-Base596,049,9201,053,184
Eos-0.8BDecision 1.0 Eos752,393,0241,053,184
Sol-2BDecision 1.0 Sol1,881,825,0882,105,856
Nox-4BQwen3.5-4B-Base4,205,751,2962,632,192
Lux-9BDecision 1.0 Lux7,936,684,5444,211,200
Vega-27BQwen3.8-27B + LoRArank-512 LoRA, 3,735,289,8565,263,872

Backbone counts are safetensors index and header sums; head counts are the decision_head.safetensors headers, read by range request.

vllm-sr/Decision-2.0-Kai-0.6B@cd49ea3 · snapshot 2026-10-06
repo size
1.52 GB
architecture
Decision2Model
task
feature-extraction
library
transformers
license
apache-2.0
safetensors
2 shards
largest file
1.50 GB
files
36
downloads
725
likes
29
decision-modelclassificationsystem-onesafetensors

Every weight of Qwen3-0.6B-Base fine-tuned plus a new 1,053,184-parameter FP32 endpoint head; the release is a uniform average of six fine-tunes (two recipes, three seeds). Raw probabilities at temperature 1, no calibration file. Apache-2.0.

repo last modified 2026-10-03

The head, from the code

The repositories ship their runtime as remote code, and decision2/_vendor/dev2model/decision_model.py is short enough to read whole (measured). Each question becomes its own prompt:

Context:
<state>
 
Task type: choice
Question:
<instructions>
Options:
<option>
{"key": "returns", "description": "Refunds, replacements and damaged deliveries"}
</option>
...
 
Select the single option best supported by the context and instructions.
Decision:

The model runs that through the backbone once. It keeps the last hidden state of each </option> block (the "candidate endpoints") and the hidden state of the final token (the "global query"). The CandidateHead scores each option as a bilinear term plus a small MLP term:

si=(Wk c~i)⋅(Wq q~)256+w⊤GELU(Wc c~i+Wq′ q~)s_i = \frac{(W_k\,\tilde c_i)\cdot(W_q\,\tilde q)}{\sqrt{256}} + w^\top \mathrm{GELU}(W_c\,\tilde c_i + W_{q'}\,\tilde q)

where c~i\tilde c_i and q~\tilde q are the layer-normed option and query states and every projection is 256 wide. A softmax over the sis_i is the answer. The head is identical across sizes; only its input width changes (1,024 for Kai up to 5,120 for Vega), which is why it grows from about one to about five million parameters.

This is the pointer family. Because the backbone is causal, an option sees the options listed before it and not those after, so option order can still move the answer; there is no reverse-and-average in the runtime. The cap is 255 options. Every question is a separate prompt, but the questions share their context prefix, which is what lets the Core ML port below run them together.

Two more things the files say (measured). config.json carries "calibration": null, and the manifest calls the output "raw probabilities (temperature 1; no calibration file)," with fixed offsets for five-level score questions; the loss has an optional Brier term, so calibration is something training may push toward, not something fitted after. And the releases are soups: Kai averages six full fine-tunes, and Vega averages LoRA adapters with their "factors concatenated along the rank," which is how a rank-512 adapter holds 3.7 billion parameters.

The 27B's card says "29.37B" parameters. That is the Qwen3.8-27B text model's 25,624,600,064 plus the 3,735,289,856-parameter LoRA plus the head, which sum to 29,365,153,792 (reasoned). Merge the LoRA, as you would to serve it, and it is a 25.6B model with a 5-million-parameter head.

What the evaluation leaves out

Line chart titled 'Jev Decision Index: balanced skill against model size, Decision 2.0 vs. Decision 1.0 and 64 public entrants', log-scale parameters on x from 0.1B to 30B, balanced skill 0 to 70 on y. A blue Decision 2.0 line rises from about 16 at 0.6B to 56.5 at Decision-2.0-Vega-27B, circled; Decision-2.0-Lux-9B is labelled 45.3. A grey Decision 1.0 line sits below. Grey dots are public entrants; a dashed Pareto frontier runs at about 57 from just under 27B to the right edge, above the Vega point. Footer: Decision 2.0 independent reproduction with the official 0.2.1 kit on the released weights; others public board snapshot 2026-09-28. Training data audited at row level against all Index test items.
Decision 2.0's own placement on the Decision Index. Vega's 56.5 is circled, and the dashed Pareto frontier passes above it: at least one public entrant at a similar size scores higher. This figure labels Lux 45.3; Lux's own card says 46.3. (vllm-sr/Decision-2.0-Vega-27B model card.)

Each card has three columns: JevArena, a "human-labelled transfer" median over 15 tasks, and the Jev Decision Index. Every size leads the same-size models it chose to list (reported): Kai 48.6 on JevArena against GLiNER2.5-Decide's 42.5, Nox 63.6 against Decider 4B's 61.9, Vega 74.0 against AutoJev-27B's 72.1. The cards are careful about the close calls — Nox and Vega are "statistically level" with the runner-up, their words.

What none of the six tables has is a row for Jev. The nearest you can get is to line up numbers from different runners on the same Index version (all reported): Vega's 56.5, Liquid's run putting Jev at 57.9 and d1 at 58.9, and the decider README placing a stock Gemma-4-31B read through its letter readout at 57.33. On that basis the largest Decision 2.0 model sits just under Jev and just under a stock checkpoint with a temperature fitted per option count (reasoned, from three separate runs). The Pareto figure says the same thing in its own picture.

Smaller things: Lux's card gives 46.3 on the Index while Vega's card puts Lux at 45.3 (measured); "training data audited at row level against all Index test items" is a contamination claim I cannot check without the data; and latencies are medians "on a single GPU," unnamed, from 4.9 ms for Kai to 71.4 ms for Vega (reported).

vllm-sr/Decision-2.0-Vega-27B@7aec49a · snapshot 2026-10-06
repo size
14.99 GB
architecture
Decision2Model
task
feature-extraction
library
transformers
license
apache-2.0
safetensors
2 shards
largest file
14.94 GB
files
35
downloads
389
likes
55
decision-modelclassificationsystem-onesafetensors

A rank-512, alpha-1024 LoRA (3,735,289,856 FP32 parameters, a uniform soup of several adapters) for Qwen3.8-27B, plus a 5,263,872-parameter endpoint head. The card's 29.37B counts the unmerged adapter. Apache-2.0.

repo last modified 2026-10-03

FluidUse: Decision 2.0 on a Mac, at 15 issues a second

FluidInference's port is the week's best engineering, and as with their GLiNER-2.5 port and SigLIP 2 port — both FluidUse, both covered — the write-up is precise about what it checked.

A frame from FluidInference's demo video. Left: a dark GitHub-style issues page for 'routerlabs / semantic-switch' with a purple 'Triage with Decision 2.0' button, a stats bar reading '787 / 1000 issues triaged, 3935 labels decided, 72 ms per issue, 13.3 issues / sec' and '5 decisions per issue in one call · Decision-2.0-Kai-0.6B · Core ML on this Mac', a sidebar of workgroup counts, and issues gaining labels such as enhancement, wg/mom-routing and priority/P2. Right: a terminal showing Apple M5 Pro usage with GPU at 97% and ANE at 0%, and per-issue log lines of about 71 ms each.
Mid-run: 787 of 1,000 issues, five labels each, at 72 ms per issue. The system monitor on the right shows where the work runs — GPU at 97%, the Neural Engine at 0%. (FluidInference, FluidUse demo video on X.)

The port. Kai and Eos become fp16 Core ML packages of 1.1 GB and 1.4 GB that run on the GPU; the Neural Engine was slower, 72 ms against 28 ms on Kai's quickstart (reported). The clever part is the packing. All questions of a request go in one call: the shared context is the prefix, each question's tokens follow, and the mask lets a question see the prefix and itself but not the others, so "every question therefore sees exactly the prompt it would see alone." Eos is harder, because its Gated DeltaNet layers carry a recurrent state rather than a cache; the port runs the prefix once and restarts each question from its final state.

Parity, measured properly. On 400 requests (2,000 decisions) from the typed-decisions test split: 0 token mismatches, 5 flipped answers for Kai and 4 for Eos, all near-ties with an upstream top-two margin of at most 0.009, and a largest probability difference of 0.0082 (reported, with the script to reproduce it). On an M5 Pro a five-question request takes 63 ms on Kai and 113 ms on Eos, against 418 and 1,595 ms for the upstream PyTorch runtime on MPS.

The demo. IssueTriageDemo takes 1,000 real issues from vllm-project/semantic-router, strips their labels, and asks five questions per issue in one call: type, owning workgroup, priority, needs-info and good-first-issue. Its README: 66.8 ms per issue, 15.0 issues a second, 5,000 decisions in 67 seconds (reported). Its check is that the labels "match the Python Core ML reference demo on every issue compared (626 of 626)" — the port agrees with itself across languages. It never compares the labels with the real ones it stripped, so the demo says nothing about whether the triage is right (measured, from the code and README).

Two details in TriageModel.swift say more (measured). The priority question's score "ranks issues well but its argmax is P0 for most issues," so the demo cuts it at this repository's 90th and 50th percentiles (1.452 and 1.251): about 10% of issues come out P0 by construction. And the yes/no flags fire at 0.6, a hand-set cut.

The one accuracy number. The parity tables are the only accuracy figures in the release, and the cards say plainly they "are not the model card's benchmark protocol." On that split, Kai answers 0.423 of choice questions, 0.628 of yes/no and 0.398 of score questions correctly; Eos 0.430, 0.598 and 0.379 (reported). A yes/no accuracy of 0.628 is not far above a coin's half (reasoned), on a held-out set that is not this model's own benchmark. It is also the only number for these two models that someone other than vllm-sr produced.

FluidInference/decision-2.0-kai-coreml@7334413 · snapshot 2026-10-06
repo size
1.21 GB
license
apache-2.0
largest file
1.19 GB
files
23
downloads
0
likes
0
coremldecision-makingon-devicesystem-one

Decision-2.0-Kai-0.6B as one fp16 Core ML multifunction package (1.1 GB, four fixed shapes), GPU-only by measurement. Parity with upstream on 2,000 decisions: 0 token mismatches, 5 near-tie flips. Apache-2.0.

repo last modified 2026-10-05

llama.cpp /v1/systemone: one endpoint, six readouts

The any model can be Jev piece found that llama.cpp had no decision endpoint, and that the nearest primitive returned the top-N tokens rather than the ones you asked about. That is now out of date. ggml-org's blog post of 2 October announces /v1/systemone in the server, "follows the System One format introduced with TypeSafe's Jev," and lists six supported models (reported): Julia-1 (144M, mmBERT-small), Laya (421M, ModernBERT-large), Kev-4B, lev (4B), OpenJev (27B, reads images) and Clef (27B), at a median 3 ms to 43 ms per question on one RTX PRO 6000.

The header graphic of the ggml-org blog post: the llama.cpp wordmark on a dark background, beside a tree of nodes in which one path from the root is drawn in orange to a single highlighted leaf among five.
The post's header: one path chosen among several. The post has no architecture diagram; the readout lives in tools/server/server-decision.cpp. (ggml-org, 'New in llama.cpp: Decision Models'.)

Its design is the opposite of SGLang's (measured, from the server code and README). SGLang turns any chat model into a letter readout with a flag. llama.cpp serves only native decision models whose GGUF carries <arch>.decision.* metadata — "a language model that classifies with prompts does not get decisions" — and it implements each family's own readout:

type in common.hhow an option is scoredused by
OPENJEVlogit of one label token per option, at the last prompt tokenOpenJev
LEVthe same; noul read from a rating scale; choice run in two orders and averagedlev
NIMBLEthe same, with every question listed in the promptBespoke Nimble
KEVdot product of the last token's hidden state with each option's end tokenKev-4B
LAYAone output column per question type, at each option's marker tokenLaya, Julia-1
CLEFall questions in one prompt; option i's score at output row iClef

Decision 2.0's bilinear-plus-MLP head has no type here, so none of this week's open family runs on llama.cpp yet (measured).

After the scores, every family goes through the same format_answer: divide by a temperature stored in the GGUF (per question type and, where the file provides them, per option-count bucket), softmax, average the variants, and report. choice confidence is TypeSafe's formula, (pmax−1/n)/(1−1/n)(p_{max} - 1/n)/(1 - 1/n); score is the expected level. The blog's example response checks against it exactly (measured): billing at 0.9049 of three options gives confidence (0.9049−1/3)/(2/3)=0.8574(0.9049 - 1/3)/(2/3) = 0.8574; the urgency probabilities 0.036, 0.1937, 0.2225 and 0.5478 give an expected level of 2.2821 and a score confidence of 0.2821. That last pair is not a coincidence: with the mode on the top of four levels, the score confidence reduces to the score minus two (reasoned, from the formula).

raw scores → probabilities · what llama.cpp's /v1/systemone does after the forward pass

as listed

1. billing
0.9049
2. shipping
0.0275
3. technical
0.0676

answer

"choice": "billing",
"probabilities": {
  "billing": 0.9049,
  "shipping": 0.0275,
  "technical": 0.0676
},
"confidence": 0.8573
At temperature 1, no bonus and one order this is the ggml-org blog's example answer. Raise the bonus with two orders on: the boost goes to billing in one order and to technical in the other, so averaging does not cancel it, it splits it between the two ends and the middle option loses.

"One forward pass" is true with qualifications, and the code is explicit about them. Questions are answered independently, each in its own prompt, except on Clef, which decides them jointly; for Kev, lev, OpenJev and Nimble the server groups the prompts so the shared state is evaluated once. And lev's choice runs twice, the second time with the options reversed, "to cancel the preference for the first label." The widget above shows what that does and does not cancel (reasoned): a first-slot bonus lands on the first option in one order and the last in the other, so averaging splits it between the two ends rather than removing it, and with three or more options the middle loses.

The choice cap depends on the model: "52 for openjev, 255 for laya and clef" (measured, README). And the server's docs end the way SGLang's did: "The probabilities are scaled with the temperatures stored in the model file. They are not guaranteed to be calibrated for your data." The blog makes the same point by example: a vague ticket "scored 0.25 with Julia-1 but 0.80 with Kev-4B." A cutoff belongs to a model, not to the endpoint.

What week four measured

Deployment is solved four ways: a closed API, Transformers remote code, Core ML, and llama.cpp — and d1, Decision 2.0's system_one() and llama.cpp all answer in the same System One JSON (reasoned).

The receipts came from the closed lab. Liquid published the file behind its chart, which items were written late, and the per-line thresholds its demo needs. Decision 2.0's comparisons stop at same-size open models. FluidUse proved its port faithful, not right. llama.cpp claimed only speed.

Calibration is still the gap. Liquid's docs call d1's outputs calibrated and publish no reliability number, and its own demo fits a threshold per line. Decision 2.0 ships raw temperature-1 probabilities. llama.cpp applies whatever temperature the GGUF carries and disclaims the result. Nobody this week shipped an ECE.

If you are choosing this week

All reasoned, from the evidence above:

The take

Week four turned decision models into infrastructure: an image-reading API, six open sizes, a native Mac runtime and a llama.cpp endpoint, all within days. The best-documented claim of the week came with its own JSON file, and checking it took ten minutes: the averages are exact, the "four of six" is generous by half a point, and the "never trained for this" inspection result needs a threshold fitted per production line.

The open releases did the expensive engineering and skipped the expensive comparison: six sizes with no Jev row, a fast demo with no accuracy, an endpoint that tells you to calibrate. All useful. None, yet, tells you how often it is right.


Sources, read on 2026-10-06: Liquid AI's d1 launch post, decision-model docs and the d1 Playground with its demos/data/compare.json and demos/data/inspection/index.json; the Decision 2.0 collection, including vllm-sr/Decision-2.0-Kai-0.6B and vllm-sr/Decision-2.0-Vega-27B, and vLLM Semantic Router; FluidInference's FluidUse (branch feat/decision-2.0 at e080acd), decision-2.0-kai-coreml and decision-2.0-eos-coreml; ggml-org's decision models in llama.cpp and llama.cpp at 43fe9c6 (tools/server/server-decision.cpp, server-context.cpp, common/common.h); and Mapika/decider with decider-2b-vision. Parameter counts are safetensors header and index reads; no third-party code was executed. Figures are reproduced for commentary from Liquid AI's blog and launch posts, the Decision-2.0-Vega-27B model card (Apache-2.0), FluidInference's demo video and the ggml-org blog; the two interactives are original.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Jev alternatives, week four: the week decision models learned where to run", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026jevalternativesweekfour,
  author = {Satyajit Ghana},
  title  = {Jev alternatives, week four: the week decision models learned where to run},
  url    = {https://ai.thesatyajit.com/articles/jev-alternatives-week-four},
  year   = {2026}
}
share