~/satyajit

Jev-Omni: a 256-way classifier with no encoders

mdjsonmcp

2026-09-22 · 16 min · explainer · llm · architecture · multimodal · benchmarks

Seven days after Jev launched, the open reproductions had covered the architecture (a System One model in 706,048 parameters), the serving trick (any model can be Jev), the on-device port (Laya on Apple silicon) and the failure mode (Jev scores zero). None of them could look at a picture.

akhilaaa3/Jev-Omni can. It takes text, an image, up to thirty seconds of audio, or sixteen frames of video, and returns a probability per option with nothing generated. The announcement calls it the first multimodal System One model and says the others miss at least a modality. That part holds up: the two nearest projects I could find are shapsider/OmniJev, whose README says audio is "wired in the client but rejected by the current backend", and OmniJev/PlayJev, a 0.8B model that reads GUI pixels and nothing else. Neither takes four modalities. This one does.

Everything else in the announcement is worth checking against the repository, and I did, because the repository publishes enough to check: a merged fp32 backbone in thirteen shards, the head as a separate head.pt, the training recipe in decision_config.json, a verification.json full of numbers nobody asked for, and two charts whose SVG text I can read the values out of.

akhilaaa3/Jev-Omni@55b53f2 · snapshot 2026-09-22
repo size
47.67 GB
task
text-classification
library
transformers
license
apache-2.0
safetensors
13 shards
largest file
4.03 GB
files
33
downloads
0
likes
10
text-classificationmultimodalmerged

Gemma 4 12B IT with a merged rank-512 LoRA and a trained 256-way head. The card's own results table puts it below Jev on the benchmark it was measured on, which is the most useful sentence on it.

repo last modified 2026-09-22

Where the decision head sits

The one thing a multimodal decision model has to explain is how four modalities reach one head. No diagram ships with the release, so here is the one I drew after reading jev_omni.py, load_model.py and both repositories' tensor lists.

four modalities, one token stream, one readout position
no encoder towers: every modality is a projection into the decoder52,379,904 parameters of stock Gemma 4 front end · 11,907,350,320 of fine-tuned decoderinputthe only weights between it and the decodertoken streamdecoder and readoutonesequencetextstate, question, numbered optionsembed_tokens · 248,320 × 3,8401 token per tokenimageone still, RGBvision_embedder · 6,912 → 3,840280 soft tokensaudiomono 16 kHz, capped at 30 sembed_audio · 640 → 3,8401 token per 640 samplesvideo16 frames, sampled by OpenCVthe image path, 16 times16 × 280 soft tokensGemma 4 decoder, 48 layers11,907,350,320 params, fp32last position onlyHead256(h − mu) / sd → Linear(3840, 256)256 logitsindex ≥ K masked to −1e30softmax over the first K
Drawn from jev_omni.py, load_model.py and the safetensors headers of both repositories. The four inlets are not four encoders: vision_embedder is one Linear over a flattened 48×48 pixel square, embed_audio is one Linear over 640 raw samples, and video is the image path run sixteen times. All three belong to stock google/gemma-4-12B-it, which the loader downloads separately; this release fine-tuned the decoder and trained the head.

The surprise is that there is nothing to draw. Gemma 4 Unified has no vision tower and no audio encoder. model.vision_embedder is a LayerNorm, one Linear(6912, 3840) and a position table — 6,912 is 48 × 48 × 3, a flattened pixel square. model.embed_audio is a single Linear(640, 3840) over raw samples. Summed from the base checkpoint's safetensors header, the entire non-text front end is 52,379,904 parameters, against 11,907,350,320 in the decoder. That is 0.44%.

receiptscaptured 2026-09-22

Jev-Omni ships 11,907,350,320 parameters of fine-tuned text decoder and 983,296 parameters of decision head. Everything that makes it multimodal — the pixel patch embedder, the vision projection, the audio projection — is 52,379,904 parameters of stock Gemma 4 that the release never touched: 0.44% of the model, downloaded from google/gemma-4-12B-it at load time.

componentparametersdtypeorigin
text decoder, 48 layers (666 tensors, 13 shards)11,907,350,320F32fine-tuned and merged in this repo
— of which the token embedding, 248,320 × 3,8401,006,632,960F32fine-tuned and merged in this repo
decision head, Linear(3840 → 256)983,296F32trained in this repo (head.pt)
vision patch embedder, 48×48 pixels → one token35,176,704BF16stock google/gemma-4-12B-it
vision → decoder projection, 3,840 × 3,84014,745,600BF16stock google/gemma-4-12B-it
audio → decoder projection, 640 × 3,8402,457,600BF16stock google/gemma-4-12B-it
everything the classifier runs, total11,960,713,520

The fine-tuned decoder has exactly the same parameter count as the stock Gemma 4 language model, which is what a merged LoRA should look like: values changed, shapes did not. The backbone is stored in fp32 — 47.6 GB — because the repository's own verification.json records that merging the adapter in lower precision moved a probability by 0.201, against 3.39e-05 in fp32.

method Read out of the artifacts, not out of the card. Each of the thirteen backbone shards and the base model's single shard was range-requested for its first eight bytes (the safetensors header length), then for the header itself, and every tensor's shape multiplied out and summed. head.pt is a torch zip archive: its four storage entries are 15,360 / 15,360 / 3,932,160 / 1,024 bytes, which at fp32 is mu[1,3840], sd[1,3840], linear.weight[256,3840] and linear.bias[256] — a Linear(3840 → 256) and two standardisation buffers.
data /articles/jev-omni/data/parameter-census.json (7 rows, 2.7 KB)

Two things follow, and they are the shape of the whole release.

First, Jev-Omni did not train the multimodal part. The repository ships a text decoder and a head. load_jev_omni downloads stock google/gemma-4-12B-it, finds its language model by walking four candidate attribute paths, and swaps in the fine-tuned decoder:

# jev_omni.py — the graft
path, old = _find_backbone(model)
parent, _, name = path.rpartition(".")
setattr(model.get_submodule(parent) if parent else model, name, decoder)
del old

The patch embedder, the vision projection and the audio projection are the ones Google shipped. Which is not a criticism — it is the correct way to build this, and it is why it took 750 optimiser steps rather than a pretraining run. It does mean the claim to be first at four modalities is a claim about the base model being first, not about the decision training being multimodal.

Second, the readout has nowhere to go but the last position. A per-option scorer needs each option encoded on its own; a marker readout like Laya's needs a [MASK] per option. Jev-Omni does neither. The prompt numbers the options, the model runs once, and a hook grabs last_hidden_state[:, -1].

The option count is in the weights

the option set is not data here — it is a slice of a trained weight matrix
"Reply with only the number of the correct option (1-K)."the prompt numbers the options; the head has one output class per number1. Yes2. No3. Not stated4. Ask a humanLinear(3840 → 256), one row per class01234567891011121314151617181920212223… up to 255K live classes, softmaxedmasked_fill(arange(256) >= counts, −1e30)class index = list position, so swapping two options scores each of them against a different trained rowthere is no row 256: `predict` raises on more than 256 options, and the card puts the supported ceiling at 20
From load_model.py (Head256, output_classes: 256 in decision_config.json) and the four storage entries inside head.pt, which at fp32 are exactly mu[1,3840], sd[1,3840], linear.weight[256,3840] and linear.bias[256]. Masking before the softmax is correct arithmetic — the unused classes contribute nothing — and it is also the thing that makes the option count a property of the weights rather than of the request.

decision_config.json says "output_classes": 256. load_model.py says what that means:

# load_model.py — the entire decision head
class Head256(torch.nn.Module):
    def __init__(self, hidden):
        super().__init__()
        self.register_buffer('mu', torch.zeros(1, hidden))
        self.register_buffer('sd', torch.ones(1, hidden))
        self.linear = torch.nn.Linear(hidden, 256, dtype=torch.float32)
 
    def forward(self, features, counts):
        z = self.linear((features.float() - self.mu) / self.sd)
        return z.masked_fill(torch.arange(256, device=z.device)[None] >= counts[:, None], -1e30)

I checked the shape rather than trusting the constructor. head.pt is a torch zip archive; its four storage entries are 15,360, 15,360, 3,932,160 and 1,024 bytes, which at fp32 is exactly mu[1,3840], sd[1,3840], linear.weight[256,3840] and linear.bias[256]. 983,296 parameters, output dimension 256.

Put that next to the sentence the CUA-S1 piece used to separate a decision head from a classifier:

This head is Linear[256, 3840]. K is welded in. The masking is arithmetically clean — unused classes contribute nothing to the softmax — and it is also the thing that makes the option count a property of the checkpoint. predict raises above 256 options, and the card is straight that quality above twenty is "not established."

This is the third independent sighting of a frozen option count in three weeks, and the first one that is not an export artefact. The browser piece found logits[batch_size, 25] in an ONNX graph. The Apple silicon piece found K frozen at 32 in five Core ML bundles and at 4 in the one that plays Snake. Both times the PyTorch model underneath still took whatever width it was handed, and the constraint arrived with the conversion. Here there is no conversion. The constraint is what was trained.

Which means the option order has to matter

The prompt is built like this:

# jev_omni.py — _prompt()
choices = "\n".join(f"{i + 1}. {value}" for i, value in enumerate(options))
return (f"{state}\n\n---\n\nQUESTION: {question}\n\nOPTIONS:\n{choices}\n\n"
        f"Reply with only the number of the correct option (1-{len(options)}).\n"
        "Output a single number and nothing else.")

Class i is list position i. Move an option from second to third and it is scored against a different row of a trained weight matrix — the same structure as openjev reading the vocabulary rows for A, B and C, which reversing the option list flipped on 27.8% of cases. The rows here were trained for the job instead of inherited from pretraining, so the flip rate could be much lower. It cannot be zero by construction, and that is the difference that matters: in a per-option scorer, order-sensitivity is not expressible; here it is a training outcome.

I could not run the test. The reference loader requires a CUDA GPU and about 50 GB of fp32 weights, and there is no GPU in the machine I write these on — the same hole the Apple silicon piece had to publish. So this is Reasoned, with a falsifier at the bottom that costs one H100-hour.

What <100ms covers

warm median latency by modality, one H200, as published
three of the four are under 100 msthe fourth is the modality the release is named forrequestmedian ms, warm100 msimageone still · ≈ 280 visual26 msaudio13 seconds, mono · ≈ 325 audio31 mstext≈ 2k tokens · ≈ 2,000 text83 msvideo16 frames · ≈ 4,480 visual504 msmedians over 20 optimised-backend requests; preprocessing and network time are excluded, so decoding the video is not in the 504 ms eithertoken counts are derived from num_soft_tokens = 280, audio_samples_per_token = 640 at 16 kHz, and 16 frames per video
All four milliseconds are the model card's own, in its own order. The hardware is an H200, not the H100 the claim names. A video request carries sixteen times an image request's visual tokens and costs nineteen times as long, so the cost is worse than linear in tokens; the config has two candidate reasons — 40 of the 48 layers use a 1,024-token sliding window, and vision tokens attend bidirectionally — and I have not separated them.

The claim that travelled is "<100ms on 1 H100." The card publishes four numbers and they are all in the table above: 83 ms for about 2k tokens of text, 26 ms for an image, 31 ms for thirteen seconds of audio, and 504 ms for a sixteen-frame video. On an H200, not an H100.

So three of the four are under 100 ms and the fourth — the modality the release is named for having — is five times over it. A video decision model quoting a text-only latency would be the finding here; this is the softer version, which is that the aggregate claim kept the three cheap modalities and dropped the expensive one. The card itself does not do this: it prints all four in a row, in the same sentence, and adds that these are "medians over 20 optimized-backend requests; preprocessing and network time are extra." That last clause is worth reading twice. Decoding the video with OpenCV and resampling the audio through ffmpeg — which predict does inline, per request — are not in any of these numbers.

"On par with Jev on Typed-benchmarks"

The card's own chart disagrees, and I can read the numbers out of it because the release ships the SVG next to the PNG.

A scatter chart titled 'Accuracy for the price', subtitled 'DecisionBench Medium, state-macro accuracy, Jev priced per question'. API cost per state runs along a logarithmic horizontal axis from one hundredth of a cent to above one cent, and accuracy up the vertical axis from 60 to 100 percent. Jev-Omni sits at 87.57 percent and Jev 1.13 just above it at 90.48 percent, both at about five hundredths of a cent per state. GPT-5.6 Luna sits at 98.76 percent at about a tenth of a cent, Gemini 3.8 Flash at 99.12 percent at about four tenths of a cent, and Claude Sonnet 5 at 99.12 percent at about one and a half cents.
The release's own headline chart, and the reason to read it rather than the announcement: on the benchmark the author built and ran, Jev-Omni is 2.9 points below Jev, and three ordinary chat models are eleven points above both. The argument the chart is making is about the horizontal axis. (akhilaaa3/Jev-Omni, assets/medium-accuracy.png.)

87.57% against Jev's 90.48% on DecisionBench Medium, state-macro. That is not on par; it is 2.9 points behind, and the author published it as the headline figure, which is the most creditable thing in the release. The second most creditable thing is that the three chat models are on the same chart at 98.76, 99.12 and 99.12 — an eleven-point gap that the launch framing of this whole category tends to leave out.

The chart is not really about accuracy. It is a cost chart, and on cost the argument is real. Jev-Omni has no bill, so its mark uses OpenRouter's Gemma 3 12B input rate against its own recorded input tokens, and it is charted at essentially the same point as Jev, whose measured cost per state is $0.00048 against Gemini's $0.00412 and Sonnet's $0.01645 — roughly ninefold and thirty-fourfold, before the correction below. The dataset card is unusually careful about exactly this — it documents a withdrawn earlier version of the chart where Jev was priced per state while Jev-Omni was priced per question, and states the correction plainly, including that per-question pricing multiplies input tokens by 2.82× on this subset. Publishing the retraction of your own favourable chart is rarer than it should be.

Three things the comparison does not settle.

The benchmark is the author's, and so is the answer key. DecisionBench is 80 scenarios and 293 questions per subset, "generated with Claude Opus 5", with answers described in its own dataset card as the "writer-intended answer." Claude Sonnet 5 scores 99.12% on it. A benchmark whose key was written by a frontier model and on which frontier models score 99 is measuring agreement with a generator as much as it is measuring decisions. That is not a reason to discard it — every synthetic benchmark in this category has the same problem, and this one at least ships its rows — but it is the denominator.

Only the easy subset is in the model card. The dataset publishes medium and hard. hard is where the interesting number is: Jev falls to 65.26% and its calibration error triples to 0.1204. Jev-Omni's card reports Medium only, and the dataset's results.json has no Jev-Omni row at all, so there is no published Jev-Omni number on the harder half of its own benchmark.

The JevBench row has no opponent. The card reports "JevBench · matched 195 groups / 231 decisions — 86.15%" with no Jev column beside it. The published JevBench figures this site has collected are 72.1% on the hard tier and 98.6% on the standard tier, so without knowing which tier the 195 matched groups came from, 86.15% is unplaceable between them.

On the two multimodal benchmarks the card is explicit that it loses: MMAU 63.10% against Inkling's 77.20%, MVBench 53.10% against Qwen3.5-397B-A17B's 77.60%. Against models 81× and 33× its size respectively, on benchmarks nobody trained for here. The card prints the parameter counts in the same table.

The calibration number, and what it is made of

A reliability diagram titled 'Does confidence match accuracy?', subtitled 'DecisionBench Medium, closer to the dashed line is better'. Predicted confidence runs along the horizontal axis and observed accuracy up the vertical, with a dashed diagonal. Jev-Omni's curve runs below the diagonal through the low and middle bins — about 0 percent observed at 15 percent stated, 13 at 30, 48 at 52, 71 at 71 — and rejoins it at the large top-right bin near 96 percent. Jev 1.13's curve sits on or just above the diagonal throughout. The three chat models appear only as large markers in the top-right corner.
Jev-Omni's reliability curve against Jev's on the same 293 questions. The headline ECE of 0.0400 is real and it is dominated by the big dot at the top right; the small bins, where a routing threshold actually lives, sit below the line. (akhilaaa3/Jev-Omni, assets/medium-calibration.png.)

The card reports ECE 0.0400 on ten bins, which is a good number and better than Jev's 0.1204 on hard — though Jev's own Medium ECE is 0.0324, so on the subset both were measured on, Jev is still ahead. The curve is the more useful artefact, and it repeats a caveat this site has made twice before: most of the mass is in the top bin, where the model is right and says so, and a low aggregate ECE is what near-saturated accuracy implies. The bins between 0.15 and 0.55 — the band where a confidence threshold earns its keep — sit below the diagonal on a handful of answers each.

Two receipts the release did not have to publish

verification.json is four lines long and worth the whole file listing:

{"verification_examples": 12,
 "merged_vs_adapter_max_probability_difference": 0.20123280584812164,
 "fp32_merged_vs_adapter_max_probability_difference": 3.3915042877197266e-05,
 "merged_vs_adapter_argmax_match": true}

Merging a rank-512 adapter into the backbone at serving precision moved a probability by 0.20. In fp32 it moved by 3.4e-05. So the repository ships fp32 weights — 11,907,350,320 parameters × 4 bytes = 47.6 GB, which is the "about 50 GB" the card quotes — and runs bf16 autocast at inference instead. That is a real finding about merging, published by the person it inconveniences, and it is the reason this download is twice the size it looks like it should be.

The second receipt is the recipe, and it is where the announcement and the artifact part company. decision_config.json records the final stage exactly:

{"size": 24000, "rank": 512, "alpha": 512, "lr": 1e-05, "epochs": 1,
 "effective_batch": 32, "steps": 750, "world_size": 4,
 "trainable_parameters": 2099183872,
 "initialization": "FP32 merged trained v1 + trained head + fresh LoRA"}

750 × 32 = 24,000, so the arithmetic is self-consistent: one epoch over 24,000 examples on four processes. The announcement says 30,000 examples on 8×H200. The gap is presumably the earlier trained v1 rank128 stage the config names as its starting point — but that stage ships no recipe, so the 30,000 total cannot be checked from the repository, and the world size that is recorded is four.

The trainable count is the other number worth sitting with. 2,099,183,872 trainable parameters — rank 512 with α=512 across the projections of a 48-layer, 3,840-wide model is 17.6% of the backbone. That is not a light-touch adapter; it is most of a fine-tune, which is consistent with a release that then merges it and ships the whole thing.

What I would actually take from this

What would change my mind

5 claims above, and what would falsify each

  1. Jev-Omni's answers change when the option list is reordered, because class index is list position.

    The cheapest possible test, and it needs one H100-hour: take fifty DecisionBench Medium Choice questions, run each with its options in the published order and again reversed, and count argmax flips and the mean maximum change in probability. A per-option scorer is at 0.0 on both by construction; rlcd-style letter readouts flip on 27.8% of reversals. If Jev-Omni comes back at exactly 0.000 on both measures, my reading of Head256 is wrong and something order-invariant is happening that I did not find in jev_omni.py.

  2. The multimodal front end is untouched stock Gemma 4.

    sha256.json covers this repository's files, not the base model's. Download google/gemma-4-12B-it, hash model.vision_embedder.*, model.embed_vision.* and model.embed_audio.*, and compare against what load_jev_omni has in memory after the graft. They should be bit-identical, because the loader never replaces them. If they differ, the release trained more than the decoder and the parameter census above understates it.

  3. The decision head is Linear(3840, 256) — 983,296 parameters with K welded in.

    python -c "import torch; print({k: v.shape for k, v in torch.load('head.pt').items()})". Anything with a trailing 1, or a shape that depends on the request, means it is a per-option scorer and I have described the wrong architecture. I read the four storage sizes inside the zip archive instead of loading it, because there is no torch in this machine; a real load is the stronger check.

  4. 504 ms for video is prefill cost, not preprocessing.

    The card says preprocessing is excluded, so the 504 ms should be reproducible by feeding sixteen pre-decoded frames straight to predict and timing only the forward pass. If the number falls sharply when OpenCV is taken out of the loop, the published figure includes decode after all and the honest video latency is higher than 504 ms, not lower.

  5. DecisionBench's answer key is a frontier model's judgment, which is why frontier models score 99 on it.

    Have two people independently label fifty medium rows and fifty hard rows and report inter-annotator agreement against the shipped key. If humans agree with the key at 99% on medium, the key is simply correct and the chat models are simply right; if they agree at, say, 90%, then the ceiling on that chart is the generator, and every model's distance from it is partly a distance from Claude Opus 5.


Nothing here was executed. The reference loader needs a CUDA GPU and about 50 GB of fp32 weights, and there is none in the machine this was written on, so every latency and accuracy figure above is Reported — read out of the model card, the dataset card's results.json, and the text nodes of the two committed SVGs. What is Measured is the arithmetic: parameter counts summed from range-requested safetensors headers across all thirteen backbone shards and the base model's single shard, and the head's shape read from the storage sizes inside head.pt's zip archive. Repository at revision 55b53f2, dataset at 19334fe, both 2026-09-22. Companion pieces: A System One model in 706,048 parameters for the two-family split this model sits outside of, and Laya on Apple silicon for the previous two sightings of a frozen option count.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Jev-Omni: a 256-way classifier with no encoders", ai.thesatyajit.com, September 2026.

bibtex
@misc{ghana2026jevomni,
  author = {Satyajit Ghana},
  title  = {Jev-Omni: a 256-way classifier with no encoders},
  url    = {https://ai.thesatyajit.com/articles/jev-omni},
  year   = {2026}
}
share