2026-10-08 · 29 min · retrieval · long-context · sparse-attention · benchmarks · inference
Why read this
Notabletop 60%Traces each UNREAL headline to its baseline, recomputes parameters, index bytes and FLOP break-even from configs, and maps it against sparse attention.
- Original analysis
- A lasting reference
- A new technique
Inference & servingNeeds datacenter GPUsPractitioner paper
How this was scored
- Is it new?
- 2 of 3: A real new idea, method or capability
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 0 of 3: Closed, nothing to run
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 2 of 3: A reference for a year or more
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 59 of 100, ranked 265 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
A RAG system and a long-context model are both answering the same question: which few pieces of all this text matter for the question in front of me? A retriever answers it out loud, with a separate embedding model, an index and a top-k. A long-context model answers it silently, inside attention, and tends to get worse at it as the haystack grows. Production stacks end up running both, and keeping two models aligned.
UNREAL, from Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel, Ron Banner, Daniel Soudry and Boris Ginsburg (all NVIDIA; Soudry is also at the Technion), says you can make the generator itself do the explicit version. Freeze the LLM. Read its hidden states to get chunk keys. Append 64 learned tokens to the question, read the hidden states at those positions to get a query, score every chunk, keep the top few, and hand their text back to the same LLM. Fewer than 500K trainable parameters. The abstract reports HotpotQA recall going from 49.1% to 73.2% on a 21M-chunk Wikipedia index, and NoLiMa accuracy going from 1.0% to 24.83% at 128K tokens.
That second number is what sent me in. A 1.0% baseline at 128K is the kind of number that says more about the baseline than the method. So I read the paper end to end, appendices included, pulled the configs of every backbone it names from Hugging Face, and redid the arithmetic it leaves implicit. There is no code release and no weights for the trained module; I looked for both and found neither, so everything below comes from the paper, the configs and my own arithmetic.
The short version: the mechanism is simple and well-motivated, and the corpus retrieval results are strong even after you account for what the baselines did not get. The long-context headline compares against the weakest line on its own chart. And "unified" means one set of weights and one scoring rule used in two quite different pipelines, which is still a useful thing to have.

Two scales of one operation
The paper's framing is the part I would keep even if the numbers were weaker. At long-context scale, evidence selection is implicit: the model receives everything and has to suppress what is irrelevant while it generates. At corpus scale it is explicit and external: a retriever picks passages before generation starts. The operation is the same, query-conditioned selection over chunks; only the count changes, from hundreds of chunks in a prompt to millions in a corpus.
If that is right, the two should be served by one mechanism, and the natural place to put it is inside the model that will read the evidence, so chunks are ranked in the representation space of their eventual reader. And the selection should be explicit, so distractors are actually removed before generation rather than merely down-weighted.
UNREAL is the decoder-only successor to INTRA (arXiv 2605.05806), the same group's earlier paper, which retrieved from the encoder memories of an encoder-decoder model through its cross-attention. Decoder-only models have no separate encoder and no cross-attention, so both halves had to be rebuilt from things a decoder does have: its residual stream.
How a hidden state becomes a key
Start with the chunk side, because it is the simpler one. Take a chunk of tokens, run it through the frozen LLM on its own, and stop at one intermediate layer . The residual-stream states at that layer are the chunk's representation:
where is the model's hidden size. One vector per token is too much to store for 21M chunks, so the token states are split into contiguous groups and each group is mean-pooled. The paper uses : every chunk becomes seven -wide vectors, whatever its length (the Wikipedia chunks are about 141 tokens).
Which layer? The paper picks it on a dev set, motivated by prior work showing intermediate
layers carry richer semantics than the last one. The chosen layers are
of 40 for Qwen3.5-35B-A3B, 47 of 52 for Muse-Glimmer-30B and 42 of 52 for
Nemotron-3.5-Lightning-30B-A3B. I checked these against each model's config.json. In
zero-based layer indices, 27 is one of Qwen3.5-35B-A3B's ten full-attention layers
(3, 7, 11, …, 39; the other thirty are linear attention), 47 is one of Muse-Glimmer's thirteen
global-attention layers (the other 39 use a 2,048-token sliding window), and 42 is the last of
Nemotron-3.5-Lightning's six attention layers (5, 12, 19, 26, 33, 42; the rest are Mamba and
MoE blocks). So in all three the index is read just after a global-attention block, late in
the stack.
What surprised me is that it barely matters. The layer ablation for Nemotron-3.5-Lightning (Table 4) tries layers 8, 18, 30, 38, 42 and 48, and complete-evidence recall@10 on HotpotQA stays between 69.42 and 72.93 across all of them. Layer 8 gives 70.25. Whatever makes this work is not a magic layer.
The query side, and where the 500K parameters live
The query is where the learning happens. UNREAL builds a retrieval input
with three parts. is the top five BM25 chunks for the question. Then the question's own tokens. Then learned embeddings , which are not words; they are free -dimensional vectors fed in where token embeddings would go. A single further learned vector sits at the very start of the sequence as a soft prompt; it is never read out, it only nudges the frozen model into "retrieval mode".
Because the retrieval tokens come last, the causal mask lets each of them attend to the BM25 context and the question. Their residual states, read at every full-attention layer , are the raw query:
The layers are mixed with learned scalars , the 64 rows are mean-pooled into groups, and each chunk is scored with ColBERT's late-interaction MaxSim:
For each of the four query vectors, find the best of the chunk's seven vectors, and add the four maxima. The top chunks by are the selection.
Notice the asymmetry. The key comes from one layer; the query is a weighted mix of all the full-attention layers, six of them for Nemotron, ten for Qwen3.5-35B-A3B, thirteen for Muse-Glimmer. That works because a residual stream is a shared workspace that every layer adds into, so states from different depths live in roughly comparable coordinates. It also explains why the method reaches across architectures: it never touches an attention head's queries or keys, so a Mamba or Gated DeltaNet block is as readable as a softmax one. That is a quiet but real advantage over every method that scores chunks with attention internals.

The paper gives the parameter count as : 64 retrieval vectors plus one soft prompt, each wide, plus one weight per mixed layer. With the hidden sizes from the configs and one per full-attention layer, that is
- Qwen3.5-35B-A3B:
- Qwen3.5-4B:
- Nemotron-3.5-Lightning-30B-A3B:
- Muse-Glimmer-30B:
So the Nemotron module is 174,726 trainable numbers. All four are under 500K, and the largest is about of a 30B model, inside the paper's "less than ". The paper states the Nemotron count of six mixed layers directly; for the others I assumed one weight per full-attention layer, which is what "a learnable sum of the full-attention outputs" implies.
Training: contrastive, with hard negatives from BM25
The loss is multi-positive InfoNCE. For a question with oracle chunks (the chunks that contain the annotated evidence) and a pool of oracles plus hard negatives:
Every oracle should outscore everything in the pool. The interesting part is where the negatives come from. For each question, the oracle chunks themselves are used as BM25 queries against the corpus; the top ten hits are thrown away because they are usually near-duplicates of the oracle, and the next 500 become hard negatives. Within a microbatch the negatives of all questions are pooled and deduplicated, with care that one question's oracle is never another's negative for that same question. These are chunks that look lexically like the evidence and are not, which is exactly the confusion a frozen model's default similarity makes.
The data is the training splits of eight Wikipedia QA sets (Natural Questions, SQuAD v2, HotpotQA, 2WikiMultiHopQA, MuSiQue, IIRC, FEVER, HoVer), all remapped onto the 21M-chunk DPR Wikipedia-2018 corpus by a KILT-style title-and-overlap alignment. Then 15K steps at a global batch of 256 questions, AdamW with , weight decay 0.1, 100 warmup steps, learning rate or . Only , the soft prompt and get gradients. The LLM never changes, which is the whole point: the model that generates is bit-for-bit the model you downloaded.
What the ablations say is doing the work
Table 3 of the paper changes one design choice at a time on Nemotron-3.5-Lightning. The column I care about is complete-evidence recall@10 on HotpotQA, baseline 72.93:
| change | HotpotQA all-R@10 | drop |
|---|---|---|
| no BM25 context ( 5 → 0) | 56.73 | 16.20 |
| read one layer instead of mixing all six | 61.01 | 11.92 |
| 16 retrieval tokens instead of 64 | 67.83 | 5.10 |
| one query group instead of four | 68.37 | 4.56 |
| one BM25 chunk instead of five | 68.83 | 4.10 |
| one pooled vector per chunk instead of seven | 69.11 | 3.82 |
| 100 hard negatives instead of 500 | 70.59 | 2.34 |
| no soft prompt | 70.86 | 2.07 |
The largest single contributor is BM25. Without the five BM25 chunks in front of the question, the method loses 16.2 points. I read this as pseudo-relevance feedback doing the job of a first hop. A two-hop question like "which instrument does the lighthouse keeper's sister play" does not name the keeper. BM25 on "lighthouse" finds the chunk that does, the retrieval tokens attend to it, and now the query can go looking for the sister by name. The interactive below reproduces this on a toy example. The second-largest contributor is the multi-layer read-out, which is the actual new idea. The fiddly knobs (token count, groups, negatives, soft prompt) are each worth two to five points.
There is one more experiment that changes how I read the paper's title. Appendix A.2.1 asks whether a frozen model retrieves at all without training: mean-pool each chunk's attention keys and the question's attention queries per head, dot them, average over heads, rank. That is the "retrieval heads" style of evidence. On NoLiMa with Nemotron-3-Nano it does beat random, but not by much: the best layer reaches an average recall@10 of about 0.24 against roughly 0.04 for random, while trained UNREAL gets about 0.65.

So the "intrinsic capability" is a weak signal that 64 trained vectors amplify a lot. That is a fine result. It is also a reminder that the retrieval here is learned, on in-domain supervision, not discovered.
Checking the corpus numbers
On the full 21M-chunk index, the paper reports complete-evidence recall@10: a question only counts if every annotated evidence chunk is in the top ten. For multi-hop questions that is a harsh metric, which is part of why the absolute numbers look low.

Reading the bars off the paper's chart:
- HotpotQA, 49.1 → 73.2. The 49.1 is Qwen3-Embedding-4B followed by the Jina reranker, the strongest baseline. The 73.2 is UNREAL on Muse-Glimmer-30B, a 30B dense model.
- 2WikiMultiHopQA, 31.7 → 60.1. Here the 31.7 is a different baseline, LateOn (a ColBERT-style multi-vector retriever) plus reranker; Qwen3-Embedding-4B plus reranker is 30.3. The 60.1 is again Muse-Glimmer.
- MuSiQue, 8.8 → 14.4, same pattern.
- Averaged over all eight datasets, 36.5 for the best baseline against 49.6 for the best UNREAL.
Each headline picks the best baseline for that dataset and the best UNREAL backbone, which is the normal way to report it. The comparison that actually convinced me is a different one: UNREAL on Qwen3.5-4B, a 4B model, gets 69.1 on HotpotQA and 57.1 on 2WikiMultiHopQA. That is the same parameter class as the Qwen3-Embedding-4B baseline, and it still gains 20 points on HotpotQA. Model size is not the explanation.
Two things keep me from taking the margins at face value. First, UNREAL is trained on the training splits of all eight evaluation datasets, mapped onto this exact corpus, with hard negatives mined from this exact corpus. The baselines are off-the-shelf. The paper points out that Qwen3-Embedding was itself trained on HotpotQA and Natural Questions, which is true, but it was not trained on 2WikiMultiHopQA's or HoVer's distribution against Wiki-2018 hard negatives. The missing control is a baseline embedder fine-tuned on the same data and negatives; without it, some of the gap is in-domain training rather than method. Second, UNREAL's query sees five BM25 chunks and the dense baselines' queries do not. Hybrid RAG fuses BM25 and dense rankings, but it does not let the BM25 hits shape the query. The ablation says that input alone is worth 16 points on HotpotQA. Remove it and Nemotron-3.5-Lightning still gets 56.73 against the baseline's 49.1, so the method survives, with a smaller margin.
Single-hop is where the gains thin out. On Natural Questions, Qwen3-Embedding-4B plus reranker reaches 26.7 and UNREAL ranges from 24.9 (the 4B) to 30.6. That fits the story: the multi-layer, BM25-conditioned query earns its keep when the question needs a bridge.
Retrieval quality carries through to answers. With Nemotron-3.5-Lightning as the generator and only the top-5 retriever varied (Table 1), HotpotQA exact match goes from 43.3 for the best baseline to 52.6 with UNREAL, against 64.3 when the model is handed only the oracle chunks. 2WikiMultiHopQA goes from 32.5 to 42.0 and MuSiQue from 13.2 to 16.9. Qwen3.5-35B-A3B as the generator shows the same ordering (48.2 against 41.3 on HotpotQA).
Inside a 128K prompt
Now the second scale. Given a long prompt, UNREAL treats it as a small corpus. It splits the text into sentence-aligned chunks of about 141 tokens that never cross a document boundary, runs each chunk through the frozen model alone up to , scores them with the same 64 retrieval tokens and the same , keeps the top , reassembles those in their original document order, and generates from that much shorter prompt. No long-context training data is involved; the modules are the ones trained for corpus retrieval.
Two things about this are easy to miss. The chunks are encoded independently, so no chunk ever attends to another during selection; and the generation pass re-encodes the selected text from scratch. Nothing from the selection pass is reused as KV cache. This is RAG over your own prompt, done with your own model, not a form of sparse attention.

Whose 1.0%?
The 1.0% is Nemotron-3-Nano reading the full 128K-token haystack and answering. Three things put that number in context.
NoLiMa is built to be hard for exactly this: the question and the needle share no keywords, so the model has to make a latent association ("lives next to the Semper Opera" answers "who has been to Dresden"). Its own paper evaluates up to 32K for most models, where GPT-4o falls from 99.3% to 69.7% and Llama 3.3 70B to 42.7%. The 64K and 128K points here are an extension.
Nemotron-3-Nano is a small-active-parameter hybrid. Its full-context score is already about 63% at 4K, against an oracle of about 96%, and below 10% by 32K. A 1.0% at 128K is that curve continuing. It is a real measurement of a real model, and a weak reader to compare against.
The fair comparison is on the same chart: two strong external retrievers with a reranker, given the same chunks and the same generator. At 128K they reach about 16% (BGE-large plus reranker) and about 14% (Qwen3-Embedding-4B plus reranker). UNREAL's 24.83% beats them by roughly nine points, which is a solid result for a retriever with no extra model. It is not a 25x improvement.
LV-Eval tells the same story more sharply. Full context falls to 49.97 F1 at 256K words; UNREAL reaches 54.66; BGE-large plus reranker sits at about 54.6 on the chart. On HELMET's RAG subset, Qwen3.5-35B-A3B reading the full context (67.2) beats both external retrievers (66.0 and 65.9), and UNREAL edges past it at 70.3. With Nemotron-3.5-Lightning the full context scores 55.5 and UNREAL 60.2.

A few loose ends I could not close from the paper:
- Which UNREAL module ran on NoLiMa? The four trained backbones are Qwen3.5-4B,
Qwen3.5-35B-A3B, Muse-Glimmer-30B and Nemotron-3.5-Lightning. NoLiMa uses Nemotron-3-Nano, and
Section 4 says the long-context modules are "the same ones presented in Section 3". The
NVIDIA-Nemotron-3-Nano-30B-A3Bconfig has the same shape as Lightning (2,688 wide, 52 layers), so a Lightning module could be paired with a Nano generator, but the paper does not say so. - Was chosen on the test set? Figure 6 shows NoLiMa accuracy peaking around twelve chunks and collapsing at both ends. The paper does not say how the in Figure 4 was picked.
- Where does come from in a prompt? The query side needs five BM25 chunks. In the long-context setting they presumably come from BM25 over the prompt's own chunks; the paper does not spell it out. On NoLiMa, which is designed to defeat lexical matching, BM25 scores near zero on its own from 32K up, so at those lengths whatever it adds is not keyword overlap.
- Did Muse-Glimmer run past its window? Its config sets
max_position_embeddingsto 131,072, yet the LOFT-style chart plots its full-context score at 256K and 512K. Part of that collapse may be a model read beyond its trained length.
The distractor experiment in the appendix is the best argument for explicit selection. With the gold chunk always present, filling the rest of a 20-chunk prompt with chunks a retriever ranked highly drops NoLiMa to about 62, against about 77 when the filler is random and about 96 for the gold chunk alone. Near-misses hurt more than noise. A model that reads everything is reading all the near-misses.

Same weights, two pipelines
The toy below lets you poke at the mechanism. Ten invented chunks about a lighthouse keeper, each with two pooled key vectors in a five-dimensional toy space, and a two-hop question. Flip between corpus mode and in-context mode, change , and turn the BM25 context off. The vectors are made up to show the mechanism. The cost panel is not: it uses the hidden sizes from the configs, the paper's and 21M chunks, and the break-even lengths from the paper's Table 6.
Switch modes and the ranking does not change: same weights, same 64 retrieval tokens, same layer mix, same MaxSim. What changes is where the keys come from and what it costs. Turn BM25 off with top-n at 2 and the violin chunk beats the sister, because the query never learned the keeper's name. That is the toy version of the paper's largest ablation.
The ranking is identical in both modes, and that is the real unification: one set of frozen weights, one set of trained retrieval tokens and layer weights, one layer to read keys from, one scoring rule, one chunk size, trained once. You do not maintain a second model whose embedding space drifts away from your generator's.
What differs is everything around the score:
| corpus index | in-context selection | |
|---|---|---|
| chunk keys | computed once offline, stored, 7 vectors per chunk | computed per request, each chunk alone, discarded |
| BM25 context | BM25 index over the corpus | BM25 over the prompt's chunks (my reading) |
| search | exhaustive MaxSim over all 21M, no ANN | MaxSim over a few hundred to a few thousand |
| what the generator gets | top in score order | top in original document order |
| KV reuse | none | none: selected text is prefilled again |
So there are two code paths, with the same scorer in the middle. That is still a lot less than a RAG stack plus a long-context stack. But "without requiring an external retriever" needs one footnote: the dense retriever is gone, BM25 is not. The paper says as much in its limitations, which mention its reliance on a cheap BM25 initial context.
What it costs
The index
Seven -wide vectors per chunk is a big index. The paper does not give its size or its precision, so I did the arithmetic, assuming bf16:
- Nemotron-3.5-Lightning, : 36.8 KB per chunk, about 790 GB for 21M chunks.
- Qwen3.5-35B-A3B, : about 602 GB.
- Muse-Glimmer-30B, : 91 KB per chunk, about 1.96 TB.
A single 1,024-wide vector per chunk, BGE-large's size, is about 43 GB for the same corpus. A ColBERT-style index at 128 dimensions per token and about 141 tokens per chunk would be about 758 GB uncompressed, so UNREAL is in the same league as late-interaction retrievers, not embedders (I measured a real one in the pplx-embed-v2-late teardown). And the paper scores every query exhaustively against all 21M chunks: on Nemotron-3.5-Lightning, 4 query vectors times 7 chunk vectors times 21M chunks of 2,688-wide dot products is about 3.2 TFLOPs and a pass over the whole 790 GB index per question. That is fine on a GPU cluster and not how anyone would serve it; there is no corpus-scale query latency in the paper.
The ablation offers a cheaper point. With the Nemotron index drops to one vector per chunk, about 113 GB, and HotpotQA complete-evidence recall@10 goes from 72.93 to 69.11. That is still 20 points above the best baseline, at a seventh of the storage. If I were building this, that is the configuration I would start from.
Building the index is a forward pass over 3B tokens through 42 of 52 layers. With the 2.87B active matrix parameters the paper's Table 6 gives Nemotron-3.5-Lightning, that is roughly FLOPs by my count, several times what a 335M-parameter embedder would spend. It is a one-off cost.
The long-context pass
For a prompt, UNREAL does more passes but less work per pass. Full context prefills tokens through all layers with quadratic attention. UNREAL encodes chunks independently through layers (no cross-chunk attention, so linear in ), runs one pass over (5 BM25 chunks, the question and 64 retrieval tokens, 910 tokens in their setting), then prefills only selected tokens for generation. Appendix C writes the FLOP difference out in full and simplifies it to
where is the active matrix parameters, the fraction of them above , the number of global-attention layers and their width. The first term is the layers UNREAL never runs on the context; the second is the cross-chunk attention it skips; the third is the retrieval pass and the re-encoding it adds. Solving gives a break-even length. The paper's Table 6 lists 4,117, 6,321 and 13,049 tokens at for Qwen3.5-35B-A3B, Nemotron-3.5-Lightning and Muse-Glimmer, and 5,569, 8,471 and 17,437 at . I redid the quadratic from the table's own and and got the same numbers within rounding. I also checked against the configs: it implies for all three, and , and are exactly that.
FLOPs are not latency. Measured time-to-first-token on one H100 with vLLM crosses over later.

Reading the chart: at 8K UNREAL is slower, about 0.6x on the two MoE models. Around 32K it breaks even. At 256K it is about 4.4x faster on Qwen3.5-35B-A3B, 2.4x on Nemotron and 1.65x on Muse-Glimmer, probably because its 39 sliding-window layers already keep its full-context cost down. The big multipliers in the appendix (146x at 10M and 1455x at 100M for Qwen3.5-35B-A3B) are projections: the baseline was only measured to 256K, and everything past that is its measured time scaled by the FLOP model.
The benchmark setup is reasonable, with caveats the paper states itself: random weights and random token ids (it measures speed, not accuracy); chunk-encoding time computed from measured per-chunk throughput rather than timed end to end; the MaxSim step and the orchestration between stages left out. The MaxSim omission is harmless, under 0.01% of FLOPs by their count. The orchestration is not free in a real server, and it is the part I would want measured.
Where it sits among attention-based retrieval
UNREAL is easiest to understand next to the methods it resembles and is not.
Retrieval heads (Wu et al., 2404.15574) showed that a few attention heads do the copying in long-context models; HydraHead builds an architecture around keeping those heads on full attention. UNREAL's untrained baseline in the appendix is in that spirit, per-head keys and queries pooled and dotted, and it is the weak line in Figure 6. UNREAL itself ignores heads entirely.
InfiniRetri (2502.12962) uses the model's attention distribution, training-free, to decide which sentences to carry forward while reading a long input in chunks. That is the closest idea: model-internal evidence selection over a prompt. The difference is that UNREAL trains a read-out for it and reuses that read-out for a corpus.
Quest (2406.10774), RetrievalAttention (2409.10516), MoBA, NSA and MiniMax's sparse attention all select inside attention: per layer, often per head, for every query token, with the whole KV cache still resident somewhere (see SparDA for what offloading it costs). They can change their mind token by token. UNREAL decides once per question, at chunk granularity, before generation starts, and the unselected text simply never reaches the generator. That makes it cheaper and blunter: it cannot go back for a chunk it did not pick, and on evidence-dense work like summarising a whole document it has nothing to offer. The paper scopes itself to sparse-evidence tasks and says so. It cites Quest, MoBA, NSA, InfLLM and Landmark Attention, but not retrieval heads, InfiniRetri or RetrievalAttention.
Against KV compression like KVzip, the trade is similar: KVzip keeps a query-agnostic subset of the cache that works for any next question; UNREAL keeps nothing and re-selects per question. And against agent-managed context in Context Language Models, UNREAL is the single-pass, no-reasoning end of the spectrum, a component an agent could call rather than an agent. The attention field guide has the background on all of the sparse-attention family.
What I would take from it
The idea I would keep is that a frozen generator's residual stream, read with a few hundred thousand trained parameters, is a better retriever for that generator than a separate 4B embedder, at least on Wikipedia-style multi-hop questions with in-domain training. The 4B result shows the effect is not about size. The ablations show where it comes from: a query that sees BM25's first hop, and a read-out that mixes every global-attention layer.
For long prompts, the honest pitch is narrower than the abstract. If your workload is sparse-evidence questions over 32K+ prompts and you already serve one LLM, UNREAL-style selection should beat both full-context reading and an external retriever by a few points, and cost less time-to-first-token from about 32K. It will not turn a 1% model into a 25% model in any sense that matters beyond the chart, because a plain retriever-plus-reranker already gets most of the way.
For a corpus, the index economics are the hard part. The seven-vector index is roughly 790 GB on Nemotron and searched exhaustively; the one-vector variant is about 113 GB and loses under four points. Neither the code nor the trained modules are released, so none of this can be tried yet.
How I checked
I read the arXiv HTML of 2610.08463v1 in full, including all four appendices and every table,
and rendered the paper's SVG figures to read values off the charts. Numbers given as "about"
are read from charts; numbers given exactly come from the paper's text or tables. I pulled
config.json for Qwen/Qwen3.5-35B-A3B, Qwen/Qwen3.5-4B,
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, meta-models/Muse-Glimmer-30B and
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 from Hugging Face and located the global-attention
layers in each. The parameter counts, index sizes, MaxSim cost and index-build FLOPs are my
arithmetic from those configs and the paper's constants, with bf16 storage as my assumption. I
re-solved the paper's Equation 10 from its Table 6 inputs. NoLiMa's evaluated lengths and its
GPT-4o and Llama 3.3 70B numbers are from the NoLiMa paper. I searched GitHub, Hugging Face and
the arXiv listing for code or weights and found none; the Hugging Face paper page lists no
linked models or datasets. I did not run anything.