2026-10-07 · 25 min · calibration · multimodal · on-device · llama-cpp · benchmarks · liquid-ai
Why read this
Notabletop 60%Reads both Open d1 checkpoints to the parameter: one reuses its LM head, one grows a new head; the 'first under 10B' board was already superseded.
- Checked against the source
- Runs on a consumer GPU
- Widely used
LLM architectureCustom licencePractitioner model
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 3 of 3: Reproduces the headline result, or shows from primary files it is wrong
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 2 of 3: A widely used model, tool or lab release
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 67 of 100, ranked 145 of 454 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
Liquid's d1 has been on this site before, as a black box. Week four covered the hosted version: a closed API that speaks Jev's wire format and reads images, with nothing published about what sits inside. On 7 October Liquid released two open-weight members of the family, d1-3B and d1-omni-600M, and I wanted the inside.
What I found is that the two models share a name, a question schema and a system_one() call, and almost nothing else. One is a vision-language model that answers through the output layer it already had. The other is a bidirectional encoder that grew a new head for the job, and it is the one that hears audio. They are built so differently that they cost different amounts for the same request, in a way the launch material never mentions.
The launch makes four claims worth checking. d1-3B "ranks first among models under 10B on the Decision Index v0.2.1". It answers a single question in 8 ms on an RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on an AGX Orin 64 GB and 50 ms on an Orin Nano. Both models have "day-one llama.cpp support". And the weights are yours to "download, fine-tune, and deploy without restrictions". The first holds with a caveat about which board, the second is Liquid's measurement and I could not rerun it, and the last two did not hold when I read the files.
One schema, two machines
Both models take the same request. A state (text, any JSON, or nothing when a picture is the whole state) plus a dict of named questions, each one of three types from Jev's schema: noul (yes or no, returns P(yes)), choice (one of named options) and score (two to ten ordered levels, returns the expected level). The answer comes back with "output_tokens": 0, because nothing is decoded. If that vocabulary is new, the llama.cpp piece walks a request from JSON to number, and any model can be Jev explains why reading a probability off an ordinary model's logits is most of the trick.
Liquid's own diagrams put the difference in the right-hand column.


The arrows tell you most of it. d1-3B reads one position per question. d1-omni reads one position per option. I went to the code to see what happens between those arrows.
d1-3B is LFM2.5-VL-3B, read at the option letters
I started with the weights. Summing every tensor in d1-3B's safetensors header gives 3,123,483,888 parameters, all bf16. LFM2.5-VL-3B, the base the card names, has exactly 3,123,483,888 too. There is no added head and no new tensor; the 707 tensors are the base model's names and shapes. Split by prefix, 2,697,198,592 are the language model (30 LFM2 blocks, 8 of them attention and 22 short convolutions, plus a 128,000-row embedding that doubles as the output layer), 412,649,712 are the SigLIP2 vision tower and 13,635,584 the projector between them.
Then I compared bytes. For the vision tower and projector, every tensor I pulled by range request, from small norms to the first 64 KB of a layer-13 attention matrix, was byte-identical to LFM2.5-VL-3B. The language model is a different story. Its norms and convolution taps moved by 3 to 6%, and the samples of its large matrices by 29 to 47% of their norm, which is far more than a fine-tune nudge. Agr moved Gemma's projections by 0.1 to 0.7%. The blog explains it: Liquid "averaged the weights of LFM2.5-2.6B and the text backbone of LFM2.5-VL-3B to create a better base model", then fine-tuned several seeds and data mixes and merged them again. So d1-3B's text half is a new blend of two Liquid models, and its eyes are LFM2.5-VL-3B's, unchanged.
The readout has no head because it uses the one that was there. prompt.py renders each question with its options under codes, "Reply with the option code only.", and the answer is the next-token distribution at the last position, cut down to the codes:
# LiquidAI/d1-3B, prompt.py:180-191
def readout(tokenizer, q: Question, logz, calibration=None) -> list[float]:
"""Option probabilities from the log-probabilities at the answer slot.
...
"""
scores = [max(float(logz[i]) for i in g) for g in readout_ids(tokenizer, q)]
if calibration is not None:
scores = calibration.apply(q, scores)
m = max(scores)
exps = [math.exp(s - m) for s in scores]
return [e / sum(exps) for e in exps]readout_ids decides which tokens count (prompt.py:155-177). A noul is the best of yes, Yes and YES against the best of no, No and NO. A choice is one letter per option, A to Z, with the space-prefixed form A pooled in, falling back to two-digit codes past 26 options. A score is the digits 0 to 9, which is why a score stops at ten levels. Each option takes its best spelling, then a softmax over the options.
Two things in that function matter. First, this is the "openjev" readout that llama.cpp already implements for other checkpoints, and that any model can be Jev traced back to SGLang's scoring endpoint. d1-3B's novelty is the training, not the readout. Second, look at calibration. The card calls the answers "calibrated", but the D1Model wrapper builds its engine as SystemOne(model=self.eval(), tokenizer=...) (modeling_d1.py:22), and calibration defaults to None. No temperature is applied, and none ships in config.json or in the GGUF's metadata. Whatever calibration d1-3B has, it learned in training. Liquid publishes no calibration error for it, and since it is not on the Decision Index board, the board has not measured one either.
The state is read once, whatever the number of questions
The part of d1-3B that is genuinely its own is hybrid.py, and its docstring says it plainly:
LiquidAI/d1-3B, hybrid.py:4-13 (module docstring)
A model runs a batch of right-padded chains `(B, T)`, or a tree: a trunk of P
tokens (the state) and N branches (the questions) that each continue it, packed
back to back on one axis with no padding, `[trunk | branch 0 | branch 1 | ...]`.
Everything that works token by token (norms, projections, the MLP) runs once
over the packed axis. A convolution continues each branch from the trunk's last
inputs. Attention runs by parts: the trunk causally, every branch token over the
trunk in one flash call that reads the trunk's keys once for all questions
(under sliding attention, one fp32 product over the trunk's last window-1 keys),
each branch over its own tokens in one varlen flash call, the two merged by
their log-sum-exp. Nothing is cached, copied or padded.LFM2 makes this harder than it sounds. Twenty-two of its 30 blocks are short convolutions with three taps, so every branch's first two tokens need the trunk's last two inputs. causal_conv gathers them with an index (hybrid.py:100-110). The 8 attention blocks split each branch token's attention into two parts, over the trunk and over its own branch, and _merge recombines them by log-sum-exp (hybrid.py:180-184), which is exact. No question sees another.
This is the design AgentJev and Agr also use: one prefix, private branches. The difference from llama.cpp's /v1/systemone is the one that piece measured: there, a causal model shares the state only while a free slot can take a copy. Here the sharing is in the attention kernel, so it holds at any question count. Liquid's own latency table shows it. On the AGX Thor, three questions over one state take 20 ms against 16 ms for one.
d1-omni-600M grows a head
The omni model is a different machine, and its encoder.py docstring is a fair summary: the trunk is LFM2.5-Encoder-350M, "LFM2 blocks (10 short convolutions, 6 GQA attention layers) made bidirectional", and the head "adds a question-type embedding, runs two pre-norm transformer layers over the text positions, and scores the hidden state at each option marker with one shared MLP."
The header agrees to the parameter. The model is 587,161,089 parameters, all float32, and the card's breakdown is exact once you group by prefix: a 354,483,968-parameter trunk (LFM2.5-Encoder-350M's count, exactly), a 26,248,193-parameter decision head, which together make the card's "381M", a 94,234,880-parameter vision encoder and projector, and a 112,194,048-parameter audio encoder with its adapter. The "600M" in the name rounds 587M up.
Each question becomes one sequence, built in prompt.py:
<bos> <state> state <q> instructions <opt> <mask> option_0 </opt> <opt> <mask> option_1 </opt> ... <decide>and the head reads every <mask>:
# LiquidAI/d1-omni-600M, encoder.py:154-167
class DecisionHead(nn.Module):
def __init__(self, d: int, layers: int):
super().__init__()
self.type_emb = nn.Embedding(3, d)
layer = nn.TransformerEncoderLayer(d, d // 64, 4 * d, 0.0, batch_first=True, norm_first=True)
self.head = nn.TransformerEncoder(layer, layers, enable_nested_tensor=False)
self.scorer = nn.Sequential(nn.LayerNorm(d), nn.Linear(d, d), nn.GELU(), nn.Linear(d, 1))
def forward(self, h, pad, marker_pos, marker_mask, qtype):
h = h + self.type_emb(qtype)[:, None, :]
for layer in self.head.layers:
h = layer(h, src_key_padding_mask=~pad)
g = torch.gather(h, 1, marker_pos[:, :, None].expand(-1, -1, h.size(-1)))
return self.scorer(g).squeeze(-1).float().masked_fill(~marker_mask, -1e4)This is the Julia-1 and Laya family: a score per mask token in front of each option, rather than a letter's logit. Because the encoder is bidirectional, every option's marker can read every other option, which the letter readout also allows and the private-branch designs forbid. The type embedding is a small, sensible touch: before the two head layers run, every position is told whether the question is a yes/no, a pick or a rating.
Unlike d1-3B, omni does ship temperatures. config.json holds ten, one per type and option-count bucket; six are fitted away from 1, from 1.175 for a 6-to-10-option choice to 1.747 for a two-option choice, with 1.666 for a yes/no. They apply to text only. The model code sets calibrate to False for any request with an image or audio (modeling_d1.py:94-98), and the card says so: "image and audio answers are the model's softmax as trained."
How media gets in without reading the question
Images and audio enter as a prefix of embeddings in front of the text, and two lines in Trunk.forward keep that prefix a function of the media alone:
# LiquidAI/d1-omni-600M, encoder.py:134-140
media_query = t < prefix[:, None]
text_key = t >= prefix[:, None]
...
mask = mask.masked_fill((media_query[:, :, None] & text_key[:, None, :])[:, None], neg)
keep_right = (t != prefix[:, None] - 1).to(h.dtype)Media positions cannot attend to text, and the centred three-tap convolution has its right tap zeroed at the last media position, so it cannot peek either. Text positions read everything. The vision path is LFM2-VL's: tiles of 512 px, 16 px patches, a 2x2 pixel unshuffle, so a 384 px picture becomes 144 positions. The audio path is a 17-layer FastConformer at one position per 80 ms, cut at 30 seconds, so at most 375 positions. A request carries images or audio, never both; passing both raises a ValueError.
The weights also show how the pieces were trained. The trunk sits about 3% from LFM2.5-Encoder-350M. The vision tower sits about 1% from LFM2.5-VL-450M's, a little more than bf16 rounding of the same weights would explain, and the projector about 12%. That fits the blog's recipe. The vision encoder was frozen while an adapter and image-only LoRA updates were trained, then "we fine-tuned the full model, merged the LoRA updates, and averaged the weights with the previous checkpoint". The tower moved only in that last step.
The cost of bidirectional
There is a price, and the launch material never states it. Because every position sees every other, the state's hidden states depend on the question after it, so they cannot be shared. probabilities_batch builds one row per question (modeling_d1.py:107-108), each carrying its own copy of the state and of the media prefix. The image or audio encoder runs once per request; the trunk runs once per question. d1-3B's trunk reads the state once. Ask omni four questions about a 3,000-token state and it reads about 12,000 positions where d1-3B reads about 3,000 plus the questions.
Two smaller limits sit in the same file. With an image, the state and question text are cut to 896 tokens (image_text_length in config.json), "as trained". And encode gives the option block a budget of max(96, min(24k + 32, max_len / 2)) tokens and cuts every option's text to an equal share without saying so (prompt.py:259-283). It is the same silent truncation the llama.cpp piece caught in Julia-1, with a roomier budget: about 31 tokens an option for three options.
The widget below follows both code paths for one request. The token counts are illustrative, but the structure is the code's.
positions the language trunk reads (one sequence)
520 positions read · the other model would read 1,320 for the same request
One tree: the state is the trunk, read once; every question is a branch that attends to the trunk and to itself, never to another question.
where the answer comes from
- logits over all 128,000 tokens at each question's last position
- keep only the option codes: yes/Yes/YES vs no/No/NO, A, B, C… or the digits 0-9
- each option scores its best spelling (max)
- softmax over the options; the shipped wrapper applies no temperature
weights in play (3123M on disk)
Two cautions from the card are worth repeating because they decide whether omni is usable. Audio was trained only on requests "between an English speaker and an assistant", so a clip of anything else is outside its training. And bf16 is unsafe: it "changed the top answer on 0.8% of text and 1.7% of audio rows", where float16 matched float32 on every check.
What "first among models under 10B" is measured on
The Decision Index is not Liquid's benchmark, and not TypeSafe's. It lives in a Hugging Face Space maintained under multimodalart, which describes itself as "unofficial and community-maintained; not affiliated with TypeSafe AI". It started as a tracker of open reproductions of Jev and grew a leaderboard. Everything it shows is in JSON files in the Space, which is what I read.
Version 0.2.1, the one d1-3B cites, is a public panel of 38 benchmarks in five areas: Knowledge (10), Language (10), Retrieval (6), Tools (5) and Arts (7). It runs 120,340 requests per model on one RTX PRO 6000. Every benchmark is chance-corrected, so 0 is guessing and 100 is perfect, and an unanswered request counts as wrong. The headline is a weighted mean of the five areas with weights 25.8%, 25.8%, 20.0%, 18.3% and 10.0%, published in the Space's data/own.json. Its board file of 28 September has 70 open entrants plus Jev, 52 of them under 10B.
d1-3B is not one of them. Both cards say so in the small print: "We scored d1-3B with the official scorer (not a leaderboard submission). All other rows come from the public leaderboard v0.2.1." I can at least check the arithmetic. Liquid's five area scores for d1-3B (23.8, 56.4, 52.8, 74.5, 36.3) under the board's weights give 48.55, against the claimed 48.57. Five inputs rounded to one decimal cannot land closer. d1-omni's areas give 15.93 against its claimed 15.95. On the 0.2.1 board, 48.57 would rank d1-3B first under 10B, above JPT-9B (46.89), and below Jev, Winnow-12B (50.02) and ten entrants of 25B parameters or more. The claim is true of the board it names.

The catch is the date. The board moved to Decision Index 0.3 on 6 October, the day before the launch. Version 0.3 changes what the headline means: the full score is 20% public benchmarks, 50% private tests of the same skills and 30% private tasks from new domains, because, in the Space's words, "most of the score comes from tests nobody can train on, so a public-only advantage counts for little." A self-run on the public kit cannot produce a 0.3 score at all.
The closest comparison is 0.3's public-only column, whose panel is 37 benchmarks instead of 38. On it, three entrants under 10B already score above 48.57: Bespoke Nimble 9B v3 at 57.19, Cloudflare's clef-flash at 56.15 and ezjev 4B s2 at 50.82. The panels differ by a benchmark, so this is not a like-for-like ranking. But "first under 10B" is a statement about the frozen 28 September board, and the field has moved since. The same goes for omni: on 0.3's public part, jiwo 0.8B scores 28.72 and Dinah-0, at 150M, scores 26.06, both well above omni's 15.95.
None of that makes 48.57 a bad number for 3B parameters. The area split is the interesting part. d1-3B's Tools score of 74.5 is the best of any model under 26B on the 0.2.1 board; JPT-9B, the next best under 10B, is at 67.0. Its Knowledge score of 23.8 is below JPT-4B's 28.7. You would expect roughly this from a small model trained hard on decisions: routing and tool choice transfer, recall of facts does not.
The other tables
The cards add internal runs on public benchmarks read as decisions. d1-3B's text table averages 77.1 over eight benchmarks; the blog's version averages 82.9 because it drops HelpSteer2, which the omni card explains "may overlap with d1-omni-600M's training data". On the seven that remain, omni averages 78.4 and leads on Civil Comments (95.8) and PAWS-X (79.5), which is the "toxicity detection and paraphrase identification" in the thread.
The vision table says d1-3B "retains the vision capabilities" of its base, 74.1 against 73.9 over eleven benchmarks. I recomputed both means and they are right, but the +0.2 is one benchmark. On VL-RewardBench d1-3B scores 65.0 against 50.9, a judging task that suits a decision model. On the other ten it averages 75.03 against the base's 76.22, and it is lower on six of the eleven, with VisualWebBench down 6.9 points and CV-Bench down 5.5. "Retains" is fair. "No cost" would not be.
How fast, and on what
Here are the latency numbers, all Liquid's. Warm calls, one request at a time, taken from the d1-3B card, which the blog repeats.
| one question | 3 questions | 3.4k-token state | 384 px image | 64 states, packed | |
|---|---|---|---|---|---|
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 / s |
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 / s |
| Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 / s |
| Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 / s |
| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms | 78 / s |
| Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms | 38 / s |
Three readings of that table change what the headline numbers mean.
The 8 ms on the 4090 is a CUDA-graph number. The card says so: model.compile(mode="reduce-overhead") runs single questions as CUDA graphs and "the RTX 4090 row uses it. Without it, a single question takes 16 ms." In runner.py only the single-question chain goes through the compiled forward (runner.py:106-110); a multi-question tree does not. So three questions on the 4090 cost 21 ms, 2.6 times one, while the blog's "three questions take only 1.3x the time of one" is true on the Jetsons, where it is 1.25x to 1.46x.
The edge numbers are cheap per question and expensive per pixel and per token. A short question fits a 30 fps frame on everything but the Orin Nano, which manages 50 ms, so 20 decisions a second. A 384 px image takes 202 ms on the Nano, about 5 a second, and a 3.4k-token state takes 1.64 s. If you are wiring a camera to a Nano, the image time decides your loop rate, not the question time.
And I could not check how the Jetson rows were run. The GPU rows say bf16, median of 20 runs. The edge rows say only that they were measured "in collaboration with NVIDIA". The card names no runtime and no precision for them. d1-3B in bf16 is 6.25 GB of weights, which is a lot for an 8 GB Orin Nano that also runs the OS. The GGUFs are 2.87 GB at Q8_0 and 1.67 GB at Q4_K_M, plus a 0.58 GB vision projector. The blog's mention of NVFP4 suggests a quantised path. The two Jetson AI Lab pages the blog links for the details returned "Page Not Found" when I fetched them. Omni has no speed numbers at all; the card says it is "an early research release".
a single question over a short text state
5 of 6 targets meet 33 ms. Vertical line: your deadline.
what has to fit in memory (d1-3B files on the Hub)
"Day-one llama.cpp support"
Liquid published GGUFs for both models the day before launch, and their cards say to run llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0 and post to /v1/systemone. So d1 goes through exactly the endpoint the llama.cpp piece took apart, which made me want to see how llama.cpp reads it.
The GGUF headers say how they expect to be read. I fetched the first 12 MB of each file and parsed the metadata. d1-3B's declares general.architecture = lfm2 and lfm2.decision.type = d1, with a systemone chat template. Omni's declares lfm2.decision.type = d1omni, lfm2.attention.causal = false, lfm2.decision.block_count = 2 (the head's two layers, folded into an 18-block model), and the same ten temperatures. Its projector file declares d1omni_v for vision and d1omni_a for audio.
llama.cpp's master branch at 18b5f8b, the head on 7 October, knows seven decision types: openjev, lev, kev, nimble, laya, clef and pplx-decider (common/common.cpp:1165-1173). d1 and d1omni are not among them. The server's server_decision_context::init throws "unsupported decision model type: " for anything else, and load_model fails on that throw. The mtmd projector list has no d1omni_v or d1omni_a, and server-decision.cpp has no audio input. I looked for the missing piece: none of the 300 most recent pull requests on ggml-org/llama.cpp mentions d1, and Liquid's fork, Liquid4All/liquid_llama.cpp, was last pushed on 17 September.
I did not build and run llama.cpp for this, so the claim rests on reading the source. But as of the commit above, a stock llama-server given these GGUFs should refuse to load them. Most likely a pull request is on its way; the GGUF metadata keys are already in gguf-py's constants, and the work is small for d1-3B, whose readout is openjev's with a different template. Until it lands, "day-one llama.cpp support" means support you cannot download from llama.cpp. The demos do not use it either: the System One Arcade Space runs transformers with --compile on an L4 GPU.
Ten demos and a robot arm

The arcade is worth reading as code, because each game is a small statement of what a decision model is for. Hand Flappy asks one question per webcam frame, "Is the person showing an open palm to the camera?", and flaps when P(yes) is at least 0.5 (flappy.js:8,43). Quick Draw is a 16-way choice. Live Triage asks six questions of a support message as you type. Nothing in it generates text. The server (game/server.py) does one thing: it cuts frames down to 512 px, calls model.probabilities() on a single worker thread, because CUDA graphs are bound to the thread that captured them, and returns the probabilities and the model time.

The Isaac Sim demo is described in the thread as d1-3B "navigating an environment", with the model on a Jetson in a hardware-in-the-loop setup. The 42-second clip shows a tabletop arm sorting blocks by colour, at four times real speed. Neither the thread nor the blog says which questions the arm asks or what plans its motion, so I can only say what it looks like: a decision model choosing the next target, with something else doing the moving.
The licence is not "without restrictions"
The blog's last section promises open weights to "download, fine-tune, and deploy without restrictions". Both repositories ship the LFM Open License v1.0. Section 5 grants commercial use only to a legal entity under a "Threshold", defined as "annual revenue of 10 million United States dollars ($10,000,000) or more"; above it, "any Commercial Use ... is not licensed under this Agreement". For most readers of this page the licence is generous; for a larger company it is a real restriction, and it is the same licence the LFM2.5 bases carry. The phrase on the blog is wrong; the cards are right.
What I would use
For anything with a camera and a fixed set of answers, d1-3B is the most interesting open decision model so far, and not mainly because of its board position. It is the only one I have read whose serving code shares the state across questions inside the attention kernel and also takes images. The three-questions-for-1.25x on a Thor is the number I would design around. The parts to check before you depend on it are the parts it leaves out: there is no temperature and no published calibration error, so fit your own thresholds on your own labelled frames. And if you plan to run it under llama.cpp, wait for the decision type to land upstream.
d1-omni-600M is what the card calls it, experimental. A 15.95 index puts it behind much smaller text models on the current board's public part, and audio was trained on one kind of English request. What it does have is the architecture I would pick for a phone-sized audio classifier: one encoder for every modality, a head that scores options in context, and temperatures fitted per question shape. It needs a version 2 more than a benchmark.
If you would rather build your own decision model than adopt one, Unsloth's training guide is the next piece on this site, at training a decision model with Unsloth.
How I checked
Sources read on 7 October 2026: Liquid's launch thread and its replies through the fxtwitter mirror; the Open d1 blog post; the LiquidAI/d1-3B and LiquidAI/d1-omni-600M cards, configs, LICENSE and every Python file (line numbers above are from those files); the cards of LiquidAI/d1-3B-GGUF and LiquidAI/d1-omni-600M-GGUF; Liquid's decision-model docs, which cover only the hosted API; the System One Arcade Space's Dockerfile, game/server.py and game scripts; and the Decision Index Space's data/index-v0.2.1.json, data/index.json (0.3), data/own.json and README.md.
Parameter counts are sums over the safetensors headers, read with HTTP range requests, for both d1 models and for LFM2.5-VL-3B, LFM2.5-Encoder-350M and LFM2.5-VL-450M. Weight comparisons are byte ranges of individual tensors: whole tensors for norms and convolution taps, the first or middle 64 KB for large matrices. The percentages are relative L2 differences over those ranges, so they describe samples, not whole layers. GGUF metadata is parsed from the first 12 MB of each file. The index arithmetic uses the area weights in own.json; fitting weights to the 71 board rows by least squares reproduces every board score to within 0.007, so the weights are the ones in use.
llama.cpp was cloned at 18b5f8b and read, not built. Pull requests were listed through the GitHub API. I did not download either d1 checkpoint or run either model, so every latency and benchmark figure on this page is Liquid's, and the "should refuse to load" above is a reading of the source, not a run.