~/satyajit

IQuest-Q1: 320B of weights, 25 layers that keep the context

mdjsonmcp

2026-10-02 · 15 min · explainer · llm · mixture-of-experts · architecture · attention · long-context · kv-cache · speculative-decoding · open-weights · agents

The interesting number in IQuest-Q1 is not 320B, and it is not 15B. It is 25.

That is how many of its 88 transformer layers keep an attention cache that grows as the conversation grows. The other 63 carry a sliding window that stops growing at 4,096 tokens. At the 512K context this model natively supports, that single structural fact decides most of what it costs to serve — and you can read it out of config.json, line by line, before anyone writes a sentence about the model.

Which is the situation you are in. IQuest-Q1 landed with open weights, a model card, a serving recipe, and a blog whose "Report" button links to #. There is no arXiv paper and no technical report — I checked the blog's own source, and the link is a published placeholder (links.report: '#'). The announcement post says "technical report / HuggingFace / GitHub — available now," but the report is the one thing that is not. So this is a read-the-config article. Everything below is grounded in the files IQuestLab actually shipped: config.json, mtp/config.json, the safetensors index, the LICENSE, the model card, and the serving flags — cross-checked against vLLM's day-0 announcement.

WeightsIQuestLab/IQuest-Q1 · Modified MIT · IQuestQ1ForCausalLM
Size320B total · ~15B active · 88 layers · hidden 3,072
Attention48 Q / 8 KV heads · head_dim 128 · 3 sliding : 1 full, sliding_window 4,096
Positionpartial RoPE rotary_dim 32 of 128 · rope_theta 1e6 (full) / 1e4 (window) · sink attention on
Experts256 routed, 8 active · intermediate_size 1,536 · layer 0 dense (12,288)
Draft headseparate mtp/ submodel · num_draft_slots 7 · served via EAGLE
Context524,288 tokens (512K) · vocab 160,000
Reportnone — the blog's "Report" link is a placeholder
IQuestLab/IQuest-Q1@5c21b06 · snapshot 2026-10-02
parameters
320.32B
repo size
650.00 GB
architecture
IQuestQ1ForCausalLM
task
text-generation
library
transformers
license
other
safetensors
176 shards
largest file
9.34 GB
files
190
downloads
604
likes
126
languages
en, zh
parameters by dtype
BF16320.32B

repo last modified 2026-09-29

320B total, ~15B active

A sparse mixture-of-experts model keeps a large pile of parameters and runs only a slice of them on each token. IQuest-Q1's slice is its 256-expert feed-forward blocks, of which a router picks 8 per token (num_experts: 256, num_experts_per_tok: 8). Everything else — attention, the embedding, the output projection, and the one dense layer at the bottom — runs every time. If you have not built one of these before, the site has a from-scratch walk through the router and the experts, and Hunyuan-A13B is the same idea at a smaller scale.

The "320B total, 15B active" on the card is worth checking rather than trusting, because both numbers fall straight out of the config. Each expert is a SwiGLU block with intermediate_size 1,536, so it holds 3 x 3072 x 1536 = 14.16M parameters (measured from the shapes). There are 256 of them in each of the 87 mixture layers — layer 0 is mlp_only, a single dense FFN of width 12,288 — which is 87 x 256 x 14.16M ≈ 315.3B just in experts. Add attention (88 x 44.04M ≈ 3.88B), the untied embedding and LM head (2 x 160000 x 3072 ≈ 0.98B), the dense layer and the routers, and the blocks sum to 320.318B (reasoned). The safetensors index reports 320,318,615,552 parameters (measured, read from the shard metadata without downloading the weights). My reconstruction lands within about a million of it — the gap is the layer norms, the attention sinks, and small biases. The config is internally honest.

Active parameters are the same arithmetic restricted to what fires on one token: attention everywhere (3.88B), 8 of the 256 experts per mixture layer (87 x 8 x 14.16M ≈ 9.85B), the dense layer and routers (~0.18B). That is ~13.9B of transformer-block compute per token; add the embedding lookup and the LM head and it reaches ~14.9B, which the card rounds to 15B (reasoned). So the headline is accurate: you store 320B and you pay for roughly 15B.

The 3-to-1 stack

Now the attention. The config does not give you a flat layer count and a single attention type; it gives you three lists that you lay end to end:

"first_layers_types":        ["full_attention"],
"hybrid_layers_types_block": ["full_attention", "sliding_attention",
                              "sliding_attention", "sliding_attention"],
"num_hybrid_layers_block":   21,
"last_layers_types":         ["full_attention", "full_attention", "full_attention"]

One full layer at the front, then the four-layer block [full, sliding, sliding, sliding] repeated 21 times, then three full layers at the back: 1 + 21*4 + 3 = 88 (measured, and num_hidden_layers is 88, so the lists are consistent). The "full" entries are the one at the front, the head of each of the 21 blocks, and the three at the back: 1 + 21 + 3 = 25. The remaining 63 are sliding-window layers with sliding_window: 4096. That is the "3 sliding : 1 full" the model card and Adina Yakup's release note both quote, counted exactly.

This is the same structural bet the site has seen twice before, and it is worth naming the family. Kimi K3's hybrid block interleaves a cheap token-mixer with an occasional expensive one; GLM-5.3-Flash runs the identical three-to-one ratio and keeps a KV cache in only eleven of its forty-five layers. IQuest-Q1 is the plainest version of the idea: the cheap layers are not linear attention or a state-space recurrence, they are ordinary attention with a short window. Nothing exotic. The win comes entirely from how the cache behaves.

88 layers · GQA 8 KV heads × 128 · bf16 · window 4,09625 of 88 layers grow a cache
25 full attention — cache grows with context63 sliding window — cache fixed at 4,096 tokens
512K tokens
IQuest-Q1 hybrid51 GiB
all-88-full baseline176 GiB
25 full layers
50 GiB
63 windowed
1.0 GiB
cache saving
3.45×

Below the 4,096-token window every layer still holds the whole context, so there is no saving. Past it, the 63 windowed layers stop growing while the 25 full layers keep climbing — the ratio walks toward its ceiling of 88/25 = 3.52×.

Why 25 of 88

A full-attention layer has to keep a key and a value vector for every token it has seen, because any future token might attend to any past one. A sliding-window layer only ever attends to the last W = 4096 tokens, so once the window is full its cache stops growing. That is the whole efficiency argument, and it is quantifiable.

IQuest-Q1 uses grouped-query attention — 48 query heads but only 8 key/value heads (num_key_value_heads: 8), each of head_dim 128 — which is the standard move to shrink the cache (the site has a field guide to attention mechanisms and a piece on pushing GQA further if the Q/KV split is new). In bf16, one layer stores, per token:

2 (K and V) x 8 KV heads x 128 x 2 bytes = 4096 bytes = 4 KiB

So the KV footprint of the whole stack at context length LL is

KV(L)=(25 L+63⋅min⁡(L,4096))×4 KiB\mathrm{KV}(L) = \bigl(25\,L + 63\cdot\min(L, 4096)\bigr)\times 4\ \mathrm{KiB}

The second term is frozen at 63 x 4096 x 4 KiB ≈ 1.0 GiB the moment the context passes 4,096 tokens. Only the first term grows. Put the model's numbers in (reasoned):

Context25 full layers63 windowedIQuest-Q1 totalall-88-fullsaving
4,0960.4 GiB1.0 GiB1.4 GiB1.4 GiB1.00×
131,072 (128K)12.5 GiB1.0 GiB13.5 GiB44 GiB3.26×
524,288 (512K)50 GiB1.0 GiB51 GiB176 GiB3.45×

At the full 512K context, the hybrid stack holds about 51 GiB of KV cache against 176 GiB for a plain 88-layer stack with the same heads — a 3.45× saving, and it keeps climbing toward the ceiling of 88 / 25 = 3.52× as the context grows. Below the window there is no saving at all: every layer still holds the whole context, so a 4K prompt costs the same either way. The saving is a long-context saving, which is the point of a model that advertises 512K.

This is exactly what vLLM's day-0 support post describes from the serving side: a hybrid KV cache coordinator "for a 3 sliding to 1 full attention layer mix, so only 25 of the 88 layers grow a cache with the context." Two independent counts — mine from the config, theirs from the allocator — land on the same 25. For the serving engine this is not just a memory number: it decides how the paged-attention block tables are laid out, because 25 layers need a growing allocation and 63 need a ring buffer. (The mechanics of that cache are in how LLM inference works.)

Holding the windowed layers together

Two smaller settings in the config exist to stop the 63 windowed layers from degrading.

Attention sinks. enable_sink_attention: true and fuse_sink_attention: true. When a sliding window slides past the start of the sequence, the softmax is forced to spend all its probability mass on a handful of recent tokens, and models trained without a safety valve push huge magnitudes into that softmax to compensate — the "attention sink" pathology that StreamingLLM first named. The fix is to give attention a permanent place to dump excess mass: a few always-visible sink tokens, or a learned per-head bias added to the softmax denominator. With 63 of 88 layers windowed, IQuest-Q1 needs this more than most; fuse_sink_attention says the sink is folded into the attention kernel rather than bolted on, and vLLM's post confirms it wired "a sinks path in the attention backends" to serve it.

Partial, split RoPE. rotary_dim: 32 with head_dim: 128 means rotary position encoding is applied to only the first 32 of each head's 128 dimensions; the other 96 carry no positional signal at all. Partial RoPE lets most of the head space stay position-agnostic, which helps a model extrapolate to lengths it did not train on. And the two layer types get different rotary bases: rope_theta: 1000000 for the full layers that must resolve positions half a million tokens apart, and swa_rope_theta: 10000 for the windowed layers that only ever span 4,096. A long base for the long-range layers, a short base for the short-range ones. The config is carrying the long-context behavior in the full layers and letting the windowed ones stay cheap and local.

The draft head: recursive MTP, served through EAGLE

IQuest-Q1 ships a second, separate model in a mtp/ folder — IQuestQ1MTP, its own config.json, its own safetensors shard. This is the multi-token-prediction head, and it is there to make decoding faster, not more accurate.

The model card states the shape in its own shorthand: "2 independent (training) / 1 recursive ×8 (inference)". In training, two independent MTP heads predict the next-next tokens and add an auxiliary loss, in the manner DeepSeek-V3 popularized. At inference, that collapses to a single head applied recursively up to 8 times: feed its output back in to draft token t+1, then t+2, and so on. The mtp/config.json backs this up with num_draft_slots: 7 (one committed token plus seven drafted ahead) and its own tiny sliding_window: 512, so the draft head is cheap to run.

Drafting is only half of speculative decoding; the main model still has to verify the draft in one parallel pass and accept the longest correct prefix. IQuest-Q1 does that verification through EAGLE, the speculative-decoding scheme that drafts from the target model's own hidden states rather than a separate small model. The serving recipe names every piece: --speculative-algorithm EAGLE, --speculative-draft-model-path $MODEL_ROOT/mtp, and the vLLM form adds "draft_sample_method": "probabilistic" — the draft tokens are sampled, not greedily argmaxed, which is what vLLM's post means by "EAGLE speculative decoding with probabilistic draft sampling ... how the recursive MTP head drafts." If multi-token prediction is new, the site's FastMTP write-up and Qwen3.8-Flash-Next both cover the training-side mechanics, and the speculative-decoding-on-vLLM piece covers the verify side. I could not read vLLM's IQuest-Q1 model file directly — it is not in the mainline tree's code index yet, and the recipe points at an IQuestLab/vllm-iquest-q1 fork I did not clone — so the serving claims here are the ones vLLM and the model card state, not ones I verified against the kernel.

How it was trained

The config tells you the shape; the blog tells you the recipe, and it is an agentic-training story more than an architecture one. Pre-training and mid-training shift the data toward code and STEM and extend the context with agentic trajectories. Then three stages build the behavior: supervised fine-tuning, reinforcement learning that produces four agentic experts (agentic user experience, multi-harness work, long-horizon tasks, and general agentic work), and MOPD — multi-teacher on-policy distillation — which folds those four experts into one student that learns on its own rollouts while the matching expert scores each token. A final cross-stage residual model merging step combines the checkpoints.

A left-to-right pipeline diagram. SFT Stage 1 feeds SFT Stage 2. From there, two paths: one initializes a student at an SFT Stage 2 checkpoint that does on-policy rollouts producing token-level advantages; the other reaches four boxed agentic experts — agentic user experience, multi-harness, long-horizon, and general. A cross-stage residual model merging step with plus-circle junctions combines the reference checkpoints, the experts, and the MOPD output into the final IQuest-Q1 model. A banner at the bottom marks the four phases: SFT, Domain SFT + RL, MOPD, Output.
The post-training pipeline: four RL agentic experts, consolidated by multi-teacher on-policy distillation, then folded in by cross-stage residual model merging. (IQuest-Q1 model card.)

One detail is more than color. The blog's own case studies show IQuest-Q1, running inside Claude Code, debugging its own training pipeline — tracing a bad reward curve to an extra space that broke prefix matching and dropped earlier turns out of the loss, and repairing a broken execution environment so infrastructure faults stopped being charged to the policy as negative reward. The model card's framing is that it "takes part in its own development under human supervision." Treat that as a description of the development loop, not a capability claim; the fixes shown are small and human-reviewed.

The benchmarks are the lab's own, and they do not lead

Here is where a read-the-config article has to be careful, because the config does not tell you whether the model is good — the benchmarks do, and these benchmarks are self-reported.

Eight bar-chart panels comparing IQuest-Q1 against other models. DeepSWE v1.1: DeepSeek-V4.1-Flash 74.2, Claude Opus 5 73.7, GLM-5.3 66.9, IQuest-Q1 64.6, Hy4-preview 64.3. NL2Repo: Claude Opus 5 75.3, IQuest-Q1 63.0, DeepSeek-V4-Pro 61.5, Hy4-preview 58.9, GLM-5.3 58.0. CyberGym: DeepSeek-V4.1-Flash 88.1, IQuest-Q1 84.5, GLM-5.3 84.5, DeepSeek-V4-Pro 83.3, Hy4-preview 78.4. Terminal-Bench 2.1: Claude Opus 5 89.1, Hy4-preview 85.4, GLM-5.3-Flash 84.3, IQuest-Q1 83.2, DeepSeek-V4-Flash 82.7. Plus JobBench, Agents' Last Exam, Humanity's Last Exam, and IQuest-CLIBench panels. IQuest-Q1's bar is highlighted in each.
IQuest-Q1 across eight coding and general-agent benchmarks, with the comparison set the lab chose. These are the lab's own evaluations; some competitor scores are their published numbers, others are the lab's own runs. (IQuest-Q1 blog.)

Read the panels honestly and IQuest-Q1 leads none of the eight. Its best showings are second place: 63.0 on NL2Repo behind Claude Opus 5's 75.3, and 84.5 on CyberGym tied with GLM-5.3 but behind DeepSeek-V4.1-Flash's 88.1. On the headline agentic-coding benchmark, DeepSWE v1.1, it scores 64.6 and comes fourth — DeepSeek-V4.1-Flash leads at 74.2 and Claude Opus 5 at 73.7. On Terminal-Bench 2.1 it is fourth at 83.2 against Opus 5's 89.1. On Humanity's Last Exam (no tools) it is third at 39.2. Most telling, on IQuest-CLIBench — the lab's own in-house benchmark for CLI experience, the one it built and could have tuned to — it scores 53.7 and still finishes behind GPT-5.6 Sol (58.5) and Claude Opus 5 (58.2).

That is not a dig; it is the right frame. IQuest-Q1 activates ~15B parameters per token and lands within a few points of frontier closed models and much larger open ones across eight agentic benchmarks. For an open-weight model you can run yourself, "a few points behind Opus 5 at a fraction of the active compute" is a real result. Just do not read the highlighted bars as wins. The lab's own benchmark notes add the fine print that matters: for each competitor it uses "the publicly reported score; otherwise, we evaluate the model" with its own harness (mini-SWE-agent for DeepSWE, Claude Code elsewhere), which means some bars are the lab's runs of rival models. There is no third-party confirmation of any of this yet, and no report describing the eval setup beyond that paragraph.

The license, and what you cannot check

The weights are Modified MIT. I read the LICENSE: it is MIT with exactly one added clause — if you use the software or a derivative in a commercial product or service, you must "prominently display 'IQuest-Q1' on the user interface." Drop the commercial-attribution clause and it is plain MIT: use, modify, distribute, sell. For most research and internal use it behaves like MIT; the string only matters if IQuest-Q1 is powering something a customer sees.

What this article could verify, it did: the parameter count (reconstructed to within a million of the shard metadata), the 88-layer 3-to-1 stack and its 25 cache-growing layers (from the config, matching vLLM's allocator count), the 4-KiB-per-token KV footprint and the 3.45× saving at 512K (arithmetic on the config), the sink and partial-RoPE settings, the MTP submodel and its EAGLE serving flags, and the absence of a technical report. What it cannot: whether the benchmark numbers hold up under an independent harness, and the exact serving-kernel behavior, which lives in a fork I did not run. When the report appears — if the # ever resolves — those are the two things to check it against.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "IQuest-Q1: 320B of weights, 25 layers that keep the context", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026iquestq1,
  author = {Satyajit Ghana},
  title  = {IQuest-Q1: 320B of weights, 25 layers that keep the context},
  url    = {https://ai.thesatyajit.com/articles/iquest-q1},
  year   = {2026}
}
share