~/satyajit

The test is whether one option can read another

mdjsonmcp

2026-09-22 · 39 min · explainer · architecture · agents · llm · calibration

You have a step in your system and you want to know whether a decision model goes there. Eighteen pieces on this site have taken that question apart in somebody else's shipped code — a browser agent, a link mapper, a code reviewer, a robot arm, a UI composer, a grep, a dinosaur game — and measured what happened. This is what the measurements add up to.

The test is four questions, and they have to be asked in this order, because the third one is the only one that cannot be bought back with money and the first two are what let you reach it.

four gates, in order — the right-hand column is what failing each one has cost somebody
supply the relation in code, then re-entergate 1 — the answer setIs the set of acceptable answers finiteand known before you make the call?noNo — then something has to write a string.379 of 385 WindTunnel tool calls carried astring property with no enum and no const:a field with no list to pick from.yesgate 2 — the enumeratorCan code you own produce that set,without a model inventing it first?noNo — then you have not removed thegenerative model, you have given it aproofreader. That is often exactly whatyou want. Price it as two models, not one.yesgate 3 — independenceCan every option be judged on its own,without reading another option?noNo — 0 of 100, below chance. Sharing acontext is not the fix: 8 of 100. Build therelation in code — intersect the identifiers,hand over one pair — then re-enter here.yesgate 4 — the thresholdIs the option list frozen, and the cutoffmeasured on your own labelled outcomes?noNo — then read the argmax and ignore thenumber. One synonym added to a dropdownmoved a winner from 0.619 to 0.215; a tunedcutoff moved 0.67 to 0.37 across two datasets.yesPut the model here.code enumerates · the model picks · code composesCost and latency are not gates. Nothing on this page was decided by them, and the two headlinemultipliers measure 4.57x and 336x per call when somebody runs both arms on one task.

Cost is not on that list, and neither is latency. That is the first thing the evidence did to my own prior, and it did it harder than I expected: when somebody runs both arms of a comparison on one task, 193.6x faster measures 4.57x, the 245x cheaper on WindTunnel's board factors into 15.25x of interface and 16.07x of price, and the 240x cheaper on a link map is 3.7x once the tuning run is counted. None of the systems below was made possible by the price. Several of them would have been fine without the model at all.

The catalogue is a good place to check the procedure against

shipwithjev.com is a hand-sorted directory of 426 builds, with its own disclaimer in the footer — "Not affiliated with TypeSafe AI. Numbers are as reported by their authors." — which is the correct disclosure and more than most such lists carry. It also sells sponsor slots at $15–20 a week, so read it as a community catalogue rather than a survey. I am using it the way a test suite is used: not as evidence that any of it works, but as a sample of what people actually reach for.

Its own category counts are the first result:

categorybuilds
Tools & apps168
Agents & browsers69
Research & data55
Games & real time47
Content & growth40
Triage & routing35
Trading & markets7
Robotics & devices5

Three of those — games and real time, trading, robotics, 59 of 426 — are the categories where this corpus has a measured ceiling. The rest are categories where the answer set arrives for free: a label set, a menu, a list of documents, a list of components. (Research & data is mostly write-ups about the model rather than builds; the directory's own type breakdown is 273 GitHub repositories, 90 X posts, 46 resources, 9 skills and 8 sites.)

That distribution is a vote for the procedure before I have argued for it, and most of the entries pass all four gates without effort. 500 emails classified for 3.5 cents, an AI slop detector at $0.00015 and 243 ms, a YouTube sponsor skipper, 700 leads scored in 40 seconds, 724 competitor ads broken down against a fixed set of facets for 9 cents, 3,282 posts answered with eight questions each for $0.1282 — every one of those is a finite label set applied per item, enumerated by the schema the author wrote, with no item's answer depending on another item's content, and the aggregation done in code afterwards. So is Stagehand on a remote browser at about $0.001 a task, which sends the accessibility tree as state and the available actions as questions, and so is typesafe-computer-use at $0.0002 per decision. These are the cases the procedure is built to say yes to, and it says yes to them.

The interesting entries are the handful where it does not, and where the reason is invisible from outside. I will take them at the gate each one lands on.

Gate 1 — is the answer set finite, and known before the call?

A decision model's output space is a list you sent in the request. It cannot write a string that is not already in it. TypeSafe says so on its own jaggedness page and the negative-space piece measured what that costs a real agent: across WindTunnel's 147 scored WebMCP attempts, 385 of 413 selected actions could not be executed until a second model wrote their arguments, and 379 of those 385 requests carried at least one string property with no enum and no const. The 28 that needed no second call were the ones the harness could run with an empty argument object.

That is not a defect, it is the type signature, and the same repository contains the artifact that shows what happens when you fight it. decision-spans.mjs turns generation into a chain of choices, exactly as the vendor's page says you can and should not: a final answer becomes eight Choice questions over the same candidate list, deduplicated and joined with "; ". That is not a sentence. It is a concatenation of spans that already existed in the prompt, and its own comment marks where it runs out.

The catalogue has one entry that lands squarely here, and it is worth naming because the failure is invisible from the outside. "Jev Calc — a smart calculator notebook ... it can calculate ANYTHING" is listed under Tools & apps. The answer set of an arithmetic expression is not finite, and the vendor's own failure-mode page lists Math and Numbers — not a calculator, does not count reliably — among its eight entries. The build that works here is the one the vendor documents: the model picks which spans are the operands, and code does the arithmetic. Its dates cookbook puts it in one line — "Extraction is a judgment, so give it to the model. Arithmetic is not, so keep it in code." I have not run the notebook and cannot say which of the two it is. I can say which of the two has a number behind it.

"Finite" is weaker than it sounds if you are running it locally

On a hosted API the option count is data. open-jev-deberta's head ends in 2.weight [1, 1024] — one scalar per option, softmax afterwards, nothing in the weights that knows how many options there are — and the cardinality caps people quote turn out to be five constants in five repositories with five unrelated causes, one of which is LETTERS = "ABCDEFGHIJKLMNOP", a string literal whose length the validator reads.

Export it and that stops being true. This corpus has found the option count welded into the exported graph twice, on two runtimes, by two unconnected authors. logits[batch_size, 25] is the graph metadata of the only open ONNX export with a browser playground: batch is symbolic, sequence is symbolic, and 25 is not. And K = 32 is frozen in five of six published Core ML bundles, with K = 4 in the sixth — the Snake bundle, whose option capacity is four because a snake has four directions to choose between. On the Neural Engine path all three axes freeze — B=1, L=96, K=32because the layout rewrite that makes the graph an ANE target is a rewrite into fixed buffers. Going past the ceiling is a re-export, not a longer array.

So gate 1, stated properly: the answer set must be finite, known before the call, and — if the model runs on the user's machine — no larger than the number somebody typed into a conversion script.

Gate 2 — can code you own enumerate it?

This is the gate people skip, and it is where the work is. It is also the one the field has already converged on without saying so: five harnesses read side by side put no model in the enumerator in four cases out of five, and the fifth added one only after a truncation dropped the control a task needed.

the same three boxes, six times — the middle one never gets bigger
ENUMERATEdeterministic code you ownDECIDEthe modelCOMPOSEdeterministic code you ownWindTunnelbrowser agentadapter.observe() — every live WebMCPtool, plus FINISH and ABSTAIN, cap 255one Choice→ an idmenu.find(id), then Mercury 2.5writes the arguments49/49 tasks, rank 1 of 21. Getting the answer set onto eight sites took 7,302 added lines across 44 tools.jev-linkmapinternal linkscandidates.py — TF-IDF cosine, headings×4, chrome filter, already-linked spans masked8,460 noulanswersplace_links — max 3 per page,one use per anchor phrase8,460 link decisions in 5.9 s for $0.27. The enumerator makes zero model calls; the cap then discards 1,704 of 2,383 yeses.jev-semgrepgrep by meaninga nested forEach — 30 lines ×however many meanings you passedone probabilityper (line, meaning)disjunctive normal form in theharness: AND and NOT are && and !309 lines of dependency-free Node. An embedding cannot do the NOT, because the line is encoded before the query exists.json-rendergenerative UIbuildCandidates — 322 lines, 42 fullyconfigured element recipes50 keys,in 2 callsstructuredClone into a tree,then validate the whole thingTwo calls for an eight-element dashboard. 668.6 bytes per candidate, so about 200 recipes fit one 32k request.Stanleycode reviewenabledHunks — identifiers a hunkdeclares, intersected with what others use23 questions,all typed69 policy constants decide whatevery returned probability means11,199 lines of source. The model's entire surface is 23 question definitions; one linked pair gets exactly one follow-up.this siterelated articleslib/related.ts — Jaccard over tag sets,then a second coverage pass— none —top four; the anchor text is thetarget's own title123 lines, zero model calls, 122 orphans to 0. The cheapest thing on this list is the one with no model in it.Box widths are semantic, not measured. The numbers under each row are.

Read the middle column. It is the same width in every row because the model's contribution really is the same size in every row: an index, or a probability. Everything else is code somebody wrote, and in every one of these repositories it is the larger half by an order of magnitude.

The clearest case is jev-linkmap, because its best idea needs no model at all. candidates.py finds, for each target page, the two-to-six word spans already present in the source page's copy that carry the target's topic terms: TF-IDF cosine with the page's own headings weighted 4x, a chrome filter that drops any page half the site already links to, and a mask over spans that are already links so nothing gets double-linked. Standard library, no dependencies, zero model calls. That is the component that turns "where should a link go" into a yes/no question, and it is the component I would build first:

# candidates.py — spans of 2-6 words already in page i's copy that speak page j's topic
sent = re.sub(r"⟦[^⟧]*?", " | ", sent)  # already a link: not an anchor
inbound = Counter(k for p in pages for k in p["links_out_all"])
# a page linked from most of the site is menu chrome, a contextual link to it adds nothing
self.chrome = {k for k, c in inbound.items() if c >= 0.5 * n}

The same shape, with a much larger bill, is what WindTunnel's 49/49 actually rests on. Getting eight self-hosted applications to expose tools took 7,302 added lines across 44 tools — about 166 lines per tool — and the leaderboard's top row is downstream of one of them. Before complete_checkout existed, all eight WebMCP configurations scored 0/3 on the Medusa store's checkout task; adding it took all eight to 3/3. Read that tool's own description and you can see it was written by somebody who knew a model would be calling it: "otherwise returns the choices so you can pass shipping_option_id" is an error path, designed in advance, that hands the model a list to pick from when it guesses wrong.

json-render states the price of admission more plainly than anyone else, and puts a number on the ceiling. Its enumerator is buildCandidates, 322 lines of hand-authored recipes for 42 configured element instances, and each candidate costs 668.6 bytes per request, dead linear, so a batched call tops out around 200 recipes at the Gateway's listed 32k context. Its README carries the consequence: "Repeating the same field in multiple forms and arbitrary new text/data are not supported." And the number in a metric reaches the model only because a human typed it twice — once into the props the renderer reads, once into an English sentence the model reads.

Nobody converged; they read the documentation

It is tempting to write that three independent teams arrived at this pattern in a week. They did not, and the truth is better: the vendor published the rule first, as a failure-mode page and four cookbooks, and these projects are following documentation. Extraction: "extract possible options using regex or a generative model and let jev-1.13 pick the correct extraction." Counting: "iterate in code over the candidates and ask one question for each, then add up the answers yourself" — which is jev-semgrep, described a week early. Re-ranking: one question per query-candidate pair, and code does the sorting.

A four-box flow diagram. A document, such as an email or invoice, feeds into a box reading regex finds candidate values, which feeds into a highlighted box reading TypeSafe picks the one that fits, annotated the judgment, which feeds into a box reading downstream code normalizes and acts on the value.
The whole procedure, as the vendor drew it before anyone built any of this. The highlighted box is the only part the model does. Reproduced here from the earlier negative-space piece, which is where this site first read it (docs.typesafe.ai, pre-parsed value extraction cookbook).

The catalogue's best entries are the ones that reproduce this without being told. "Keystroke oracle / predictive launcher — your launcher ranks by aliases, fuzzy match, and..." is jev-linkmap's architecture applied to an app launcher: the launcher already holds a finite list of items, fuzzy matching is the enumerator, and the model re-ranks what fuzzy matching found. Nothing in it can name an app you do not have.

Gate 3 — can every option be judged without reading another?

This is the gate with a number on it, and the number is the most informative measurement in the category.

On an independent fifteen-task benchmark, Jev takes thirteen tasks, scores 100% on rule and evidence judgment, 100% on ordinal scoring and 98% on ARC-Challenge — and then scores 0 of 100 on relational choice, questions that "use information in one option to select another."

Zero is not the bottom of a bad scatter. If the task were three-way and the model were guessing, the chance of scoring exactly zero is (2/3)100(2/3)^{100}, about one draw in 4×10174 \times 10^{17}; at four options it is 3.2×10133.2 \times 10^{-13}, at five 2.0×10102.0 \times 10^{-10}. There is no plausible option count that makes 0/100 a bad day. Zero out of a hundred is what you get from a deterministic rule that is orthogonal to the answer key — a scorer ranking each option on local plausibility, landing on the same wrong answer every single time.

The mechanism is one line of somebody else's source. A per-option scorer's encoder takes one option; row["options"] is never passed; the options meet for the first time in softmax(log_odds), four scalars long, after every forward pass has finished. An option that refers to another option is referring to something that is not in its context. Not down-weighted — absent.

And here is where my own prior needed narrowing. The obvious fix — use a model whose options share a context — does not work either:

Relational choice — 100 cases, accuracy %
Qwen3-4B prototype — all options in one prompt
53%
Laya 421M — all options in one sequence, bidirectional
8%
Jev 1.13 — each option scored alone
0%
050100

Laya's encoder is ModernBERT and bidirectional; every option attends to every other option, in both directions, through a shared scalar head. The options could not see each other more thoroughly. It scores 8. So option isolation is sufficient to make relational choice impossible, and shared context is necessary and nowhere near sufficient. The honest form of gate 3 is that no model in this family can follow a reference between options, and switching families does not buy you one.

Which is why failing gate 3 is work, not a verdict

Three codebases hit this and all three did the same thing: build the relation in deterministic code, then hand the model one pair and ask an independent question.

Stanley wants to know whether a hunk that looks unrelated to the task is there because another hunk needs it. It cannot ask that. So enabledHunks does it with two regexes — the identifiers each hunk declares on added lines, intersected with the identifiers each hunk references — and only then opens a frame containing exactly two hunks and asks one yes-or-no. The rationing is the instructive part: the follow-up uses .find, not .filter, so only a hunk already flagged weak, already linked to a hunk already scoring well, gets a follow-up, and it gets exactly one.

The xArm7 harness does the same thing where it is impossible to miss. A movement is a joint choice over X, Y, Z and the gripper — 81 combinations, and the right value on one axis depends on the others. The harness never asks for it. Each cycle is one Choice over eight intents, then four independent Choice questions, one per channel, three options each, with the constraint written into the instruction text: "This question controls ONLY X ... do not answer for another axis." The model returns four signs. The executor supplies every millimetre, and its comment is the design statement: it "never changes the selected sign."

jev-semgrep is the purest version. Each line, each meaning, one independent probability; then disjunctive normal form evaluated in the harness over those probabilities. Its README states the principle better than I can: "Because each meaning yields an independent probability, logical AND and NOT are plain boolean operations, not a trick with set differences or negative queries." An embedding cannot do this, because the line is encoded before the query exists. A cross-encoder judging a proposition per line can, and then && is just &&.

The catalogue has a case that sits exactly on this line and cannot be adjudicated from outside, which makes it the most useful entry in it. Post scoring with SuperX — 61 questions per draft in about a second for $0.0004 — reports that it "picks the viral post 2 in 3 times", on a model "fitted on 9,481 real posts from 207 creators." Picking between two posts is a comparison. If each post is scored alone and code compares the scalars, it passes this gate by construction and the architecture is right. If the comparison is inside the question, it is the 0/100 case, and the aggregate accuracy will hide it — reversing an option list moved 27.8% of one suite's individual answers while its accuracy barely twitched. One reversed list and an afternoon settles which.

The same gate, one scale up: the branching problem

There is an argument going around that agents do not have a reasoning problem, they have a branching problem: every action creates more possible next states, the tree explodes over long horizons, autoregressive reasoning is an expensive way to search it, and therefore "search cheap, branch wide, reason only when necessary" should get more valuable as horizons lengthen.

I think the shape of that is right and the conclusion about who does the searching is wrong, and the measurement is already in.

A forward pass produces one distribution over the caller's options and nothing else. A 255-option Choice carries at most log22557.99\log_2 255 \approx 7.99 bits — one call, eight bits, no tokens, no text, no state. There is no scratchpad, no place to write down "I am three hops in and my target is Rubber Duck" except the state the caller reassembles on the next call. So a decision model can price one branch. It cannot hold a frontier, and it cannot sequence what it prices.

What that predicts is exactly what the benchmark found: on a Wikipedia navigation task with a ten-hop budget, Jev reached the destination on 1 of 120 routes and stopped on 118 — the one task in the suite that requires carrying a plan across steps, from the model that takes thirteen of the suite's fifteen. The same argument appears verbatim in the systems built on it. Vercel's json-render lists "Root selection, grouping, and deciding when to stop require planning, which is a documented weakness" in its own README. The Jev Engineering guide states it as a batching rule — "The one rule: questions cannot read one another's answers. If a decision depends on a fresh search result, run the search first." Stanley states it as a limit of its own product: "hunks are judged individually; intent spread across unlinked hunks is not modeled."

The catalogue supplies the test case, and the contrast with a build that works is almost too neat. Under Agents & browsers there is a headless Chromium agent whose demo enters the Wikipedia page for "Café" and navigates hyperlinks until done — which is the WikiRouter task, the one with the 1/120 behind it. Under the same broad problem domain, jev-linkmap makes 8,460 link decisions in 5.9 seconds and works. Same subject matter, same primitive, opposite outcomes, and the difference is not difficulty: linkmap's 8,460 decisions are 8,460 independent questions about candidate pairs a TF-IDF pass already enumerated, and none of them needs to know what the previous answer was.

So the branching argument survives with its subject changed. Cheap node evaluation is genuinely valuable as horizons lengthen — but the enumerator has to enumerate successors, and the harness has to hold the frontier, apply the budget and decide when to stop. That is a planner you wrote, with a fast scorer bolted into it. It is a real and good thing to build. It is not the model doing the search, and the one published measurement of asking the model to do the search is 1 of 120.

Gate 4 — will code branch on the probability?

If the answer is no, take the argmax and stop reading. If the answer is yes, this gate is harsher than any of the three above it, and it is the one I had filed as secondary.

Start with the measurement that changed my mind. Take a five-option list, score it, then add paraphrases of the correct answer one at a time — synonyms a second engineer might reasonably add to a dropdown.

same evidence, same model, same correct answer — three synonyms added to the dropdown
acts on its ownp ≥ 0.5escalatesto a person0.000.250.500.751.00probability returned for the option0.6190.6195 optionsthe list as written0.5890.4826 options+ 1 paraphrase0.7620.2657 options+ 2 paraphrasesargmax is now a paraphrase0.8210.2158 options+ 3 paraphrasesargmax is now a paraphrasep(concept) — every phrasing of the right answer, summedp(the original option) — the number your threshold reads2.292.122.032.18raw logit, scale 2.0–2.3The belief is flat. The denominator grew. A 0.5 policy acts at five options and escalates at six, on identical evidence.Measured on a local per-option scorer — the architecture the hosted model is argued to share. Nobody has run it against the API.

The raw entailment logit of the original option is flat: 2.29, 2.12, 2.03, 2.18, drifting only because the padded batch width moves. The model's belief in the concept rises from 0.619 to 0.821, correctly, because the paraphrases are also right. And the number your threshold reads falls from 0.619 to 0.215, through the cutoff, on evidence that never changed. A policy that acts above 0.5 and escalates below it acts at five options and abstains at six.

This is not a bug and no amount of training fixes it. Each option is encoded in its own row and scored in isolation, then the softmax is taken over the scalars, so the probability an option receives is a function of what else the caller happened to list. The calibration you are relying on is calibrated against your option list, not against the world.

Three more findings stack on top, and together they make the threshold the least portable quantity in the system:

Two catalogue entries land here, in opposite directions. Fraud detection with Jev and Kimi K3 — 100 emails classified in 1.42 s, the uncertain ones routed to a larger model, 96 of 100 correct for about $0.07 — is the confidence-gated cascade done right, and it is worth naming because an earlier piece here reported those numbers and could not find a citable source for them. The catalogue has one; it is the author's own post. The threshold in it is still a threshold somebody picked, and the Banking77-to-Web-of-Science result says it will not survive being pointed at different email.

The other direction is Trading & markets, seven builds, including one described as "I gave Jev $10,000 and let it trade." I cannot verify anyone's deployment and I am not going to try. What I can say is what the corpus found when it read the code of the trading bot it could read: MODEL=mock is the default even when an API key is present, the dry run is the default, and no win rate, no P&L and no profitability number appears anywhere in the repository. Its own README states the mechanism and stops. That is the right place to stop, and it is a long way from a result.

What the evidence did to my own prior

I went in with a three-clause test: finite action space, enumerable by code you control, each option judged on its own merits. That survived as the primary test. Four things moved.

Cost and latency demoted further than I expected. Not just secondary — mostly not measuring the model. Local-versus-hosted comparisons measure the network on one side and a forward pass on the other, which is why the Apple silicon ports publish 1.394x about themselves and get quoted at 50x, and why the browser piece refused the same comparison in its own favour. The three-figure speed multiplier is only reachable against a baseline doing the thing the whole argument says to stop doing: at the measured 0.707 s per call, 193.6x needs the LLM side to spend 136.9 seconds on every call.

receiptscaptured 2026-09-22

Every headline multiplier in this category, put next to the number somebody got when they ran both arms of the comparison. None of them is fabricated and none of them survives its own denominator. The largest surviving gap in the whole table is 16x, and it is a price-per-token ratio rather than a property of the architecture.

the claim as it travelswhosemeasuredwhat the difference is
193.6x fasterTypeSafe homepage; relayed as "up to 200x", recirculated this week as Jev Engineering4.57xper-call latency, both arms on one xArm7 task. 193.6x needs the LLM side to spend 136.9 s per call — a reasoning trace, not a classification.
444.6x cheapersame336.0xper-call cost, same trial. Same order of magnitude; the gap is task mix. This is the one headline that broadly survives.
245x cheaperWindTunnel board, vs GPT-6 Astra on screenshots15.25x × 16.07xInterface term times price term. Hold the interface at WebMCP and the price term is all that is left; against the best conventional WebMCP row it is 2.3x.
112x cheapersame board, vs code execution6.94x × 16.07xSame 16.07x price term. The interface half is the transferable one, and it is available to any model that uses it.
240x cheaper than Opusjev-linkmap announcement3.7x$17.91 all-in against $66.37, once the $15.51 rubric loop and the $2.07 editor pass are counted. Break-even on the tuning is 32,147 pages; the site has 566.
zero pages against 566same, the race chart18xPer page, matched on the same 120 pages. A cost claim and a throughput claim drawn as one fact. At Jev's own concurrency Opus finishes the site in 3.3 minutes.
50x faster than Jevattributed to laya-coreml, which never says it1.394xThe only speed ratio either repo publishes about itself — Neural Engine against compiled MLX, same machine. The 50x divides a local forward pass into an HTTPS round trip.
63x fasterLFM2.5-350M-RLCD1.09x – 62.91xThe ratio is (95.4 + 15.21·G) / (48.6 + 0.0392·B): it is how many tokens you made the baseline generate. The 62.91x case generated 226; the 1.09x case generated 13.
roughly 60 decisions per secondlaya-coreml, via a third-party summary49.66/s60 is the row in the repo's own sweep that misses its deadline on 100% of ticks. A Snake decision is also three calls carrying 4, 2 and 2 options, not one.
at most 1 GBsame summary1,833 MiB943.6 MiB is peak MLX allocation for one short question. Ten full-context questions is 1.79 GiB. Neither is process RSS, and the weights are unquantised FP16.
99.7% against Jev's 83.6%cua-s1-formsin-domain vs coldA 706K-parameter specialist measured on data from the generator it trained against, versus a generalist that has never seen the distribution. The release says so itself.
only a mild hit in accuracysaid of jeff, a frozen-encoder server37.8% vs 72.1%JevBench hard tier, using JevBench's own runner. 0 of 7 on ambiguous items, 1 of 19 on long policies. Latency and cost claims hold; the third one does not.
45% to 65% recall for $15.51jev-linkmap's System 2 loop64.2%, for $0Moving two thresholds on data already collected captures 91% of the gain, at higher precision. The two rubric rewrites add the last 9%.

The pattern is not that anyone lied. In every case the claimed ratio is a real division of two real numbers; what moves is which two. Three failure shapes recur: a hosted round trip divided by a local forward pass (the Atlantic, not the model), a bounded question compared against a model writing a reasoning trace (a baseline the argument itself says to stop using), and a recurring cost quoted without the one-time cost that made it work.

method Each 'measured' figure is the one computed in this site's own teardown of that claim, from the artifact the claimant published: openroboto-ai/jev-robot-control's seed-0 pair for the speed and cost multipliers, nekuda-ai/WindTunnel's canonical results.csv for the interface/price factorisation, stas4000/jev-linkmap's out/ and runs/ for the link-map costs, mizorewww/laya-coreml's own benchmark JSON for the Apple silicon ratios, notnotsamuel/LFM2.5-350M-RLCD's committed cases for the speedup regression. Each row links back to the piece that did the arithmetic.
data /articles/where-to-use-jev/data/multipliers.json (13 rows, 5.6 KB)

Calibration promoted from a consideration to a gate. See above. It is the second structural question, not a tuning detail.

"Finite" needs "at export time" bolted on, if the model runs on the user's machine. Three exports, three frozen option counts.

And the fourth is the one that changes how much any published comparison is worth. An independent team spent several days trying to reproduce, break and improve this architecture, and reports that the same checkpoint scores 76.9% in-distribution and 54.1% out of it. I have no attributed source for it — the attribution attempt is its own piece, and it failed — so treat it as Reported, but it is the third independent measurement of the same gap in this corpus, and the other two agree with it:

who measured itin-distributionout-of-distributiongap
typed-decisions, DeBERTa-v3-large, baseline85.2%62.2%23.0
typed-decisions, same, with augmentation84.6%64.8%19.8
the unattributed reproduction above (Reported)76.9%54.1%22.8

A fourth, from a different angle: kev's out-of-domain Brier score is 0.339 against the hosted model's 0.211, on 764 records from six public sources it never trained on — the only out-of-domain comparison in the whole batch, and the project that ran it published its own loss.

The out-of-distribution split that matters is not a held-out row, it is a held-out question: a new instruction with a new option set over the same states. Splitting states tells you almost nothing, because the model can memorise "this slot, this option list, this wording."

That number reframes every head-to-head this site has covered. cua-s1-forms' 99.7% against Jev's 83.6% is a 706K-parameter specialist measured on data from the generator it trained against, against a generalist measured cold — the release says so itself. Laya's 0.766 against 0.727 comes from the checkpoint fine-tuned on that benchmark's own training split, which its own README prints above the fold, next to base checkpoints that score near chance zero-shot. A twenty-point in-to-out gap means "X beats Jev on typed decisions" is, mostly, a claim about X's training set. It is still the correct way to demonstrate that a specialist beats a generalist on its own turf, which is worth demonstrating. It is not the sentence it gets quoted as.

The release that makes the point best is the one that broke the pattern. Bespoke's Nimble measured itself on its own 324-row contrastive holdout — its own data, its own generator, every advantage — and scored 90.12% against Jev's 93.21%. Then published it, at the top of the chart, with the raw counts. That is what an in-distribution comparison looks like when it is run honestly, and it is the only one in the batch that went the other way.

The same reproduction reports two more things that bear on where you spend effort. Six architectural ideas in one night — anchor selection, continuity smoothing, asymmetric windows, bucketed temperature, two-stage retrieval, full RLCD — all failed or failed to generalise, and moving state outside the problem sequence to save compute cost 27 points, apparently by discarding the masked-language-model priors the encoder was pretrained with. The only thing that worked was data: 1,200 to 123,475 examples took out-of-distribution accuracy from 40.69% to 54.98%.

And then the result that should be read next to any model can be Jev:

Out-of-distribution accuracy, same evaluation — Reported, source unattributed
trained encoder, 1,200 examples
40.69%
a second trained encoder
54.11%
same encoder, 123,475 examples
54.98%
a third trained encoder
56.28%
a decoder, nothing trained
56.71%
0204060

A stock decoder, read through its own option logits with nothing trained, beat every trained small encoder in that comparison. That is the same result openjev got from a different direction — an untrained 4B reaching 84.5% modal agreement against Jev's published 88.3% on the vendor's own released cases — and the same result pngwn found on Gradio pull requests, where the untouched base model read through letter logits was as good on the judgement questions and better above 2,000 tokens. The readout is a serving feature, not a model feature, and it has been sitting in SGLang's /v1/score since June 2025, built for rerankers, fifteen months before the category existed.

Which sharpens gate 2 into something practical: before you buy a decision model, find out how much of what you want is the readout. The readout is free. What is not free is the enumerator, and — per the reproduction above — the data.

Five places I would not put one, even though it would work

The procedure tells you when a decision model can go somewhere. These are the cases where it can and I would not.

1 · You are already making a generative call at that step. WindTunnel is the sharpest instance and the numbers are all in its own transcripts. Mercury runs on 0.93 calls per step under WebMCP — almost every tool call needs arguments — so the decision model is a second round trip inside a step that already contains a model. It buys 22.4% of the bill: Jev is $0.045 of the $0.202 spent across 147 attempts, and the argument writer is the other 77.6%. It costs latency: Jev's median call is 403 ms, Mercury's is 614 ms, and about 63% of the attempt's 3.21 s median agent time is model latency.

What it buys back is containment, and that is a real product: menu.find(id) throws if the returned id was not offered, so the action space is the runtime's and not the model's. If your actions are reversible and cheap, ask whether you need that. If one of them charges a card, you probably do.

2 · Code already decides. The Jev Engineering guide's own sorting rule says it: "An exact rule, such as stopping after ten actions, belongs in code, not in either model." Stanley's eight policy blocks hold 69 hard-coded constants — 41 bare probability thresholds, two aliased to a shared one, 25 caps and a minimum count — and those constants, not the model, decide what every answer means. And the sharpest version is this site's own: lib/related.ts is 123 lines of tag arithmetic, zero model calls, and it took the orphan count from 122 to 0. The thing a decision model would have bought — a link placed on the right words — is real and I do not have it. The thing it would not have bought is the coverage, which was the actual problem.

3 · The corpus is fixed and the query set is open — index it instead. jev-semgrep is a beautiful idea with the wrong unit for a search box: one pass over this site's prose is 2.7M input tokens, eleven cents and a hard floor of 83 seconds, charged to you rather than the person typing, on every query, because there is no index. This site's /api/search is BM25 over contextualised chunks, built once at startup, returning in single-digit milliseconds for free. Replacing it would be a straight downgrade. The version that survives is the flipped loop: O(queries × lines) becomes O(propositions × lines) if you pick a small fixed set of propositions and score the corpus against them at build time. Eleven cents per proposition per build, zero per query, no model on the request path.

4 · A threshold sweep on data you already have would do. jev-linkmap's System 2 loop rewrote its rubric twice for $15.51 and reports recall going 45% to 65%. Holding the hand-written rubric fixed and moving two thresholds to the tuned values scores 64.2% recall at 84.0% precision — higher precision than the trained rubric achieves, for nothing. The threshold move captures 91% of the total gain. That is a free axis on data you already collected, and it is worth sweeping before you pay for anything else, including a model.

5 · The loop is real time and the model is hosted. This one is a timing constraint rather than a design one, and it is the only item here where the fix is mechanical. A hosted model in a real-time loop is bounded by the round trip, not by the model. The T-Rex harness measured the takeoff window at 13 to 26 frames, median 18, against a 22-frame round trip at 370 ms — so 88.5% of jump-or-duck obstacles offer at most one chance, and nothing guarantees that chance lands inside the window. The fix is not a faster model; you cannot make the window wider or the network shorter. It is concurrency: four questions per round trip turns a cadence wider than 88.5% of the windows into one narrower than all of them. The robot arm is the same arithmetic with the same answer — 0.62 Hz against the 3.125 Hz the simulation's own clock needs, with 87.9% of the wall time spent queueing on two sequential calls that could have been one request.

That is the honest reading of the catalogue's 47 games-and-real-time builds: the turn-based ones have no window problem at all — one entry reports 0.7 s per move on Slay the Spire 2, which is free time in a card game — and the frame-locked ones are bounded by a number that has been measured and has a known fix.

The tier below the tier

One more thing the evidence says, and it comes from outside this family entirely.

Uber's ADR runs the same two-tier discipline on 10,000+ agent sessions a day: a cheap triage pass on everything, an expensive investigating agent only on what it flags. Its triage layer resolves 40.7% of tasks at $0.017 and 2.3 s; the escalated path costs $0.029 and 29.7 s — 1.7x the money and 12.9x the wall clock. (That article's prose attributes the 13x to cost; it is the latency ratio. My arithmetic, correcting my own site.) Removing the triage layer raises recall from 0.667 to 0.805 and costs 29% more while introducing false positives, so the cheap tier is load-bearing for precision, not only for the bill.

Two things about it are worth taking. The triage model is an ordinary LLM, not a decision model — the tiering discipline does not require this architecture, it requires a cheap pass with an escalate-on-ambiguity bias. And the layer underneath both of them, the one that catches the most, is regex and entropy with no model at all, at 97.2% precision across 212 unique credentials. ADR's own sentence is the boundary: simple non-LLM checks work well for known static patterns "but fail for attacks requiring reasoning about tool semantics, causal context, and enterprise policies."

A three-column comparison of one apple-to-plate task. Jev 1.13, GPT-6 Astra and GPT-4.1 mini each show a final simulation frame, the final tool-centre-point coordinates, the chosen intent with a probability, and four separate columns for X, Y, Z and fingers, each with its own probability. Jev's are 98, 96, 93 and 86 percent and labelled native Jev probabilities; both GPT columns read 100 percent across the board and are labelled self-reported JSON probabilities. A sidebar gives cumulative cost of $0.0188, $5.9336 and $0.2885, decision cycles of 113, 106 and 160, and wall times of 181.8, 707.3 and 704.3 seconds.
Why the tiers are not a line you slide along. The frontier model placed the apple in seven fewer cycles than the decision model and cost 336x more per call to do it; the cheap LLM — the one a spectrum reading puts just above the decision tier — spent 15.3x the decision model's total and failed the task outright. One seed, one trial, per the repository's own caveat (openroboto-ai/jev-robot-control, media/final.png).
openroboto-ai/jev-robot-control@7a4ed8b · snapshot 2026-09-22
tracked files
88
license
MIT
branch
main
tests
1 file
source
178.0 kB
commit date
2026-09-19
source by language
Python140.7 kB(21)HTML37.3 kB(2)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-09-22 at 7a4ed8b branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, testFileCount

Where the corpus disagrees with itself

Four things I could not reconcile, which a reader applying this procedure should know are open.

Option-order instability is either 0.13 or 0.00, and both numbers are in this corpus. A strict per-option scorer must score exactly zero here: reversing the caller's array changes nothing about any individual (state, question, option) triple. The negative-space piece called 0.13 — quoted without a citation in a third party's BENCHMARKS.md — "the single most load-bearing unverified figure in this whole category." A day later, a third party measured it at 0.00% across 36 permutations. Thirty-six items is thin, and neither piece notices the other. If 0.13 is real, strict isolation is wrong and the likeliest explanation is the documented two-stage path for high-cardinality choices.

There are four expected calibration errors for the same model, spanning an order of magnitude: 0.0313 on a 1,200-item MMLU probe, 0.246 in the figures circulating from a competitor's chart, 0.077 on kev's run, and a range of 0.038 to 0.261 across thirteen subsets in Bespoke's public suite. No article reconciles them, and they probably cannot be reconciled, which is itself the finding: "calibrated" is not one thing you can be.

There is no such thing as this model's latency. Published p50s in this corpus run 111 ms, 178 ms, 233 ms, 236–276 ms, 246.7 ms, 338.6 ms, 351 ms, 383 ms, 403 ms and 707 ms — a 6x spread for the same model, and every article reaches for whichever one its arithmetic needs. They are not inconsistent; they are different request shapes over different network paths. The load-bearing gap is that nobody has timed a large request: the 37-question, 161-option, ~7,100-token composition call that json-render's "rendered in milliseconds" claim depends on has never been measured at all.

And one internal typo worth naming so it does not propagate. The three-tiers piece counts Stanley's policy constants as 69 in its body — with the breakdown, 41 + 2 + 25 + 1 — and as 64 twice afterwards. 69 is the one consistent with its own arithmetic.

What would change my mind

6 claims above, and what would falsify each

  1. The test is gate 3 — independence between options — and it is the only one that cannot be bought back with money.

    The decisive experiment is a single architectural change, and nobody has run it: take a per-option scorer and let the options attend to each other, changing nothing else. If relational choice, order stability and cardinality all move together, the gates are not independent and this procedure is over-factored. If letting options attend fixes relational choice and introduces order sensitivity — which is what I predict from the 53% and the 8% — then gate 3 is a fork with two prices and no model pays both. If a model ships that solves relational choice at the same latency and cost, gate 3 stops being structural and this whole article is a snapshot of one generation of models.

  2. The 0/100 on relational choice and the 'questions cannot read one another's answers' batching rule are the same constraint at two scales.

    They may be two different things: one a rule about parallel questions within a request, the other a failure within a single question's option set. If a probe shows the model handling cross-option reference fine inside one question while failing only across questions in a batch, then the batching rule is about request parallelism, gate 3 is fixable by issuing sequential requests, and the "same gate, one scale up" section is an over-reading. This is one afternoon and an API key, and it changes the procedure.

  3. The reported probability is a function of the caller's option list, so a confidence threshold is not portable.

    Measured on a local per-option scorer — DeBERTa-v3-xsmall through transformers.js — not on the hosted model, which nobody outside the vendor can inspect. Duplicate an option against the hosted API and score again. If the winning probability does not fall by roughly what the duplicate takes, the hosted model is doing something after the per-option softmax that nobody has described, and that would be the most interesting undocumented thing about it. Until somebody runs it, gate 4 rests on an architectural argument plus a local reproduction.

  4. Every published head-to-head in this category is in-distribution for the challenger and cold for the incumbent, and the gap is worth about twenty points.

    The 76.9% / 54.1% pair is Reported and I have no attributed source for it; the 85.2% / 62.2% and 84.6% / 64.8% pairs come from one project's committed run reports, on one corpus, with a three-seed spread of 2.5 points on the out-of-domain number alone — so anything smaller than that is noise. Cross-evaluate any two of the open reproductions on each other's held-out sets — Nimble's contrastive holdout, cua-s1's real-form eval and Laya's typed-decisions set are all published — and report both directions. If the challengers hold up off their own generators, this claim is too strong and the head-to-heads mean more than I am saying.

  5. The enumerator is the hard part and the cheapest component.

    Counted by reading source, not by running anything: 7,302 lines of tool code for WindTunnel's eight sites, 322 lines of recipes for json-render's 42 candidates, 11,199 lines of Stanley against 23 questions, and a TF-IDF candidate generator that makes no model calls at all. The way to falsify it is to derive an enumerator instead of writing one — component candidates from prop schemas plus a data source, WebMCP tools from an existing OpenAPI surface, anchor candidates from an embedding index. If a released project generates its option sets rather than hand-authoring them, the most laborious part of this goes away and gate 2 becomes a formality.

  6. A stock decoder read through its own option logits matches a trained small encoder out of distribution.

    Reported, from an unattributed reproduction, and it agrees with two things this site measured independently — an untrained 4B at 84.5% modal agreement against a published 88.3%, and a base model read through letter logits matching its own fine-tune above 2,000 tokens. It is falsified by any careful comparison on a held-out question split where the trained encoder wins by more than seed noise. Note what it would not falsify: that the readout is free. That part is a handler in SGLang and has been since June 2025.

The one-paragraph version

Put a decision model where the answer set is finite before you call, where code you own can enumerate it without a model inventing it first, and where every option can be judged on its own merits. The third clause is the one with a measurement on it — 0 of 100, below chance, and switching to a model whose options share a context gets you 8 — and failing it is work rather than a verdict: build the relation in deterministic code, hand over one pair, ask one independent question. The same constraint one scale up is why planning lives in the harness and why a Wikipedia navigation task scores 1 of 120 while 8,460 independent link decisions take 5.9 seconds. If code is going to branch on the probability, freeze the option list and measure the cutoff on your own outcomes, because adding a synonym to a dropdown moved a winner from 0.619 to 0.215 and a tuned threshold moved 0.67 to 0.37 between two datasets. Everything else you have been told is a ratio with a denominator worth checking: 193.6x is 4.57x, 245x is 15x of interface times 16x of price, 240x is 3.7x, and the readout that makes any of it possible has been a serving flag since June 2025. The model is the smallest, cheapest and least interesting component in every system here. The enumerator is the product.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "The test is whether one option can read another", ai.thesatyajit.com, September 2026.

bibtex
@misc{ghana2026wheretousejev,
  author = {Satyajit Ghana},
  title  = {The test is whether one option can read another},
  url    = {https://ai.thesatyajit.com/articles/where-to-use-jev},
  year   = {2026}
}
share