2026-10-02 · 15 min · explainer · llm · mixture-of-experts · architecture · attention · long-context · kv-cache · speculative-decoding · open-weights · agents
The interesting number in IQuest-Q1 is not 320B, and it is not 15B. It is 25.
That is how many of its 88 transformer layers keep an attention cache that grows as the conversation grows. The other 63 carry a sliding window that stops growing at 4,096 tokens. At the 512K context this model natively supports, that single structural fact decides most of what it costs to serve — and you can read it out of config.json, line by line, before anyone writes a sentence about the model.
Which is the situation you are in. IQuest-Q1 landed with open weights, a model card, a serving recipe, and a blog whose "Report" button links to #. There is no arXiv paper and no technical report — I checked the blog's own source, and the link is a published placeholder (links.report: '#'). The announcement post says "technical report / HuggingFace / GitHub — available now," but the report is the one thing that is not. So this is a read-the-config article. Everything below is grounded in the files IQuestLab actually shipped: config.json, mtp/config.json, the safetensors index, the LICENSE, the model card, and the serving flags — cross-checked against vLLM's day-0 announcement.
| Weights | IQuestLab/IQuest-Q1 · Modified MIT · IQuestQ1ForCausalLM |
| Size | 320B total · ~15B active · 88 layers · hidden 3,072 |
| Attention | 48 Q / 8 KV heads · head_dim 128 · 3 sliding : 1 full, sliding_window 4,096 |
| Position | partial RoPE rotary_dim 32 of 128 · rope_theta 1e6 (full) / 1e4 (window) · sink attention on |
| Experts | 256 routed, 8 active · intermediate_size 1,536 · layer 0 dense (12,288) |
| Draft head | separate mtp/ submodel · num_draft_slots 7 · served via EAGLE |
| Context | 524,288 tokens (512K) · vocab 160,000 |
| Report | none — the blog's "Report" link is a placeholder |
- architecture
- IQuestQ1ForCausalLM
- task
- text-generation
- library
- transformers
- license
- other
- safetensors
- 176 shards
- largest file
- 9.34 GB
- files
- 190
- downloads
- 604
- likes
- 126
- languages
- en, zh
repo last modified 2026-09-29
320B total, ~15B active
A sparse mixture-of-experts model keeps a large pile of parameters and runs only a slice of them on each token. IQuest-Q1's slice is its 256-expert feed-forward blocks, of which a router picks 8 per token (num_experts: 256, num_experts_per_tok: 8). Everything else — attention, the embedding, the output projection, and the one dense layer at the bottom — runs every time. If you have not built one of these before, the site has a from-scratch walk through the router and the experts, and Hunyuan-A13B is the same idea at a smaller scale.
The "320B total, 15B active" on the card is worth checking rather than trusting, because both numbers fall straight out of the config. Each expert is a SwiGLU block with intermediate_size 1,536, so it holds 3 x 3072 x 1536 = 14.16M parameters (measured from the shapes). There are 256 of them in each of the 87 mixture layers — layer 0 is mlp_only, a single dense FFN of width 12,288 — which is 87 x 256 x 14.16M ≈ 315.3B just in experts. Add attention (88 x 44.04M ≈ 3.88B), the untied embedding and LM head (2 x 160000 x 3072 ≈ 0.98B), the dense layer and the routers, and the blocks sum to 320.318B (reasoned). The safetensors index reports 320,318,615,552 parameters (measured, read from the shard metadata without downloading the weights). My reconstruction lands within about a million of it — the gap is the layer norms, the attention sinks, and small biases. The config is internally honest.
Active parameters are the same arithmetic restricted to what fires on one token: attention everywhere (3.88B), 8 of the 256 experts per mixture layer (87 x 8 x 14.16M ≈ 9.85B), the dense layer and routers (~0.18B). That is ~13.9B of transformer-block compute per token; add the embedding lookup and the LM head and it reaches ~14.9B, which the card rounds to 15B (reasoned). So the headline is accurate: you store 320B and you pay for roughly 15B.
The 3-to-1 stack
Now the attention. The config does not give you a flat layer count and a single attention type; it gives you three lists that you lay end to end:
"first_layers_types": ["full_attention"],
"hybrid_layers_types_block": ["full_attention", "sliding_attention",
"sliding_attention", "sliding_attention"],
"num_hybrid_layers_block": 21,
"last_layers_types": ["full_attention", "full_attention", "full_attention"]One full layer at the front, then the four-layer block [full, sliding, sliding, sliding] repeated 21 times, then three full layers at the back: 1 + 21*4 + 3 = 88 (measured, and num_hidden_layers is 88, so the lists are consistent). The "full" entries are the one at the front, the head of each of the 21 blocks, and the three at the back: 1 + 21 + 3 = 25. The remaining 63 are sliding-window layers with sliding_window: 4096. That is the "3 sliding : 1 full" the model card and Adina Yakup's release note both quote, counted exactly.
This is the same structural bet the site has seen twice before, and it is worth naming the family. Kimi K3's hybrid block interleaves a cheap token-mixer with an occasional expensive one; GLM-5.3-Flash runs the identical three-to-one ratio and keeps a KV cache in only eleven of its forty-five layers. IQuest-Q1 is the plainest version of the idea: the cheap layers are not linear attention or a state-space recurrence, they are ordinary attention with a short window. Nothing exotic. The win comes entirely from how the cache behaves.
Below the 4,096-token window every layer still holds the whole context, so there is no saving. Past it, the 63 windowed layers stop growing while the 25 full layers keep climbing — the ratio walks toward its ceiling of 88/25 = 3.52×.
Why 25 of 88
A full-attention layer has to keep a key and a value vector for every token it has seen, because any future token might attend to any past one. A sliding-window layer only ever attends to the last W = 4096 tokens, so once the window is full its cache stops growing. That is the whole efficiency argument, and it is quantifiable.
IQuest-Q1 uses grouped-query attention — 48 query heads but only 8 key/value heads (num_key_value_heads: 8), each of head_dim 128 — which is the standard move to shrink the cache (the site has a field guide to attention mechanisms and a piece on pushing GQA further if the Q/KV split is new). In bf16, one layer stores, per token:
2 (K and V) x 8 KV heads x 128 x 2 bytes = 4096 bytes = 4 KiB
So the KV footprint of the whole stack at context length is
The second term is frozen at 63 x 4096 x 4 KiB ≈ 1.0 GiB the moment the context passes 4,096 tokens. Only the first term grows. Put the model's numbers in (reasoned):
| Context | 25 full layers | 63 windowed | IQuest-Q1 total | all-88-full | saving |
|---|---|---|---|---|---|
| 4,096 | 0.4 GiB | 1.0 GiB | 1.4 GiB | 1.4 GiB | 1.00× |
| 131,072 (128K) | 12.5 GiB | 1.0 GiB | 13.5 GiB | 44 GiB | 3.26× |
| 524,288 (512K) | 50 GiB | 1.0 GiB | 51 GiB | 176 GiB | 3.45× |
At the full 512K context, the hybrid stack holds about 51 GiB of KV cache against 176 GiB for a plain 88-layer stack with the same heads — a 3.45× saving, and it keeps climbing toward the ceiling of 88 / 25 = 3.52× as the context grows. Below the window there is no saving at all: every layer still holds the whole context, so a 4K prompt costs the same either way. The saving is a long-context saving, which is the point of a model that advertises 512K.
This is exactly what vLLM's day-0 support post describes from the serving side: a hybrid KV cache coordinator "for a 3 sliding to 1 full attention layer mix, so only 25 of the 88 layers grow a cache with the context." Two independent counts — mine from the config, theirs from the allocator — land on the same 25. For the serving engine this is not just a memory number: it decides how the paged-attention block tables are laid out, because 25 layers need a growing allocation and 63 need a ring buffer. (The mechanics of that cache are in how LLM inference works.)
Holding the windowed layers together
Two smaller settings in the config exist to stop the 63 windowed layers from degrading.
Attention sinks. enable_sink_attention: true and fuse_sink_attention: true. When a sliding window slides past the start of the sequence, the softmax is forced to spend all its probability mass on a handful of recent tokens, and models trained without a safety valve push huge magnitudes into that softmax to compensate — the "attention sink" pathology that StreamingLLM first named. The fix is to give attention a permanent place to dump excess mass: a few always-visible sink tokens, or a learned per-head bias added to the softmax denominator. With 63 of 88 layers windowed, IQuest-Q1 needs this more than most; fuse_sink_attention says the sink is folded into the attention kernel rather than bolted on, and vLLM's post confirms it wired "a sinks path in the attention backends" to serve it.
Partial, split RoPE. rotary_dim: 32 with head_dim: 128 means rotary position encoding is applied to only the first 32 of each head's 128 dimensions; the other 96 carry no positional signal at all. Partial RoPE lets most of the head space stay position-agnostic, which helps a model extrapolate to lengths it did not train on. And the two layer types get different rotary bases: rope_theta: 1000000 for the full layers that must resolve positions half a million tokens apart, and swa_rope_theta: 10000 for the windowed layers that only ever span 4,096. A long base for the long-range layers, a short base for the short-range ones. The config is carrying the long-context behavior in the full layers and letting the windowed ones stay cheap and local.
The draft head: recursive MTP, served through EAGLE
IQuest-Q1 ships a second, separate model in a mtp/ folder — IQuestQ1MTP, its own config.json, its own safetensors shard. This is the multi-token-prediction head, and it is there to make decoding faster, not more accurate.
The model card states the shape in its own shorthand: "2 independent (training) / 1 recursive ×8 (inference)". In training, two independent MTP heads predict the next-next tokens and add an auxiliary loss, in the manner DeepSeek-V3 popularized. At inference, that collapses to a single head applied recursively up to 8 times: feed its output back in to draft token t+1, then t+2, and so on. The mtp/config.json backs this up with num_draft_slots: 7 (one committed token plus seven drafted ahead) and its own tiny sliding_window: 512, so the draft head is cheap to run.
Drafting is only half of speculative decoding; the main model still has to verify the draft in one parallel pass and accept the longest correct prefix. IQuest-Q1 does that verification through EAGLE, the speculative-decoding scheme that drafts from the target model's own hidden states rather than a separate small model. The serving recipe names every piece: --speculative-algorithm EAGLE, --speculative-draft-model-path $MODEL_ROOT/mtp, and the vLLM form adds "draft_sample_method": "probabilistic" — the draft tokens are sampled, not greedily argmaxed, which is what vLLM's post means by "EAGLE speculative decoding with probabilistic draft sampling ... how the recursive MTP head drafts." If multi-token prediction is new, the site's FastMTP write-up and Qwen3.8-Flash-Next both cover the training-side mechanics, and the speculative-decoding-on-vLLM piece covers the verify side. I could not read vLLM's IQuest-Q1 model file directly — it is not in the mainline tree's code index yet, and the recipe points at an IQuestLab/vllm-iquest-q1 fork I did not clone — so the serving claims here are the ones vLLM and the model card state, not ones I verified against the kernel.
How it was trained
The config tells you the shape; the blog tells you the recipe, and it is an agentic-training story more than an architecture one. Pre-training and mid-training shift the data toward code and STEM and extend the context with agentic trajectories. Then three stages build the behavior: supervised fine-tuning, reinforcement learning that produces four agentic experts (agentic user experience, multi-harness work, long-horizon tasks, and general agentic work), and MOPD — multi-teacher on-policy distillation — which folds those four experts into one student that learns on its own rollouts while the matching expert scores each token. A final cross-stage residual model merging step combines the checkpoints.

One detail is more than color. The blog's own case studies show IQuest-Q1, running inside Claude Code, debugging its own training pipeline — tracing a bad reward curve to an extra space that broke prefix matching and dropped earlier turns out of the loss, and repairing a broken execution environment so infrastructure faults stopped being charged to the policy as negative reward. The model card's framing is that it "takes part in its own development under human supervision." Treat that as a description of the development loop, not a capability claim; the fixes shown are small and human-reviewed.
The benchmarks are the lab's own, and they do not lead
Here is where a read-the-config article has to be careful, because the config does not tell you whether the model is good — the benchmarks do, and these benchmarks are self-reported.

Read the panels honestly and IQuest-Q1 leads none of the eight. Its best showings are second place: 63.0 on NL2Repo behind Claude Opus 5's 75.3, and 84.5 on CyberGym tied with GLM-5.3 but behind DeepSeek-V4.1-Flash's 88.1. On the headline agentic-coding benchmark, DeepSWE v1.1, it scores 64.6 and comes fourth — DeepSeek-V4.1-Flash leads at 74.2 and Claude Opus 5 at 73.7. On Terminal-Bench 2.1 it is fourth at 83.2 against Opus 5's 89.1. On Humanity's Last Exam (no tools) it is third at 39.2. Most telling, on IQuest-CLIBench — the lab's own in-house benchmark for CLI experience, the one it built and could have tuned to — it scores 53.7 and still finishes behind GPT-5.6 Sol (58.5) and Claude Opus 5 (58.2).
That is not a dig; it is the right frame. IQuest-Q1 activates ~15B parameters per token and lands within a few points of frontier closed models and much larger open ones across eight agentic benchmarks. For an open-weight model you can run yourself, "a few points behind Opus 5 at a fraction of the active compute" is a real result. Just do not read the highlighted bars as wins. The lab's own benchmark notes add the fine print that matters: for each competitor it uses "the publicly reported score; otherwise, we evaluate the model" with its own harness (mini-SWE-agent for DeepSWE, Claude Code elsewhere), which means some bars are the lab's runs of rival models. There is no third-party confirmation of any of this yet, and no report describing the eval setup beyond that paragraph.
The license, and what you cannot check
The weights are Modified MIT. I read the LICENSE: it is MIT with exactly one added clause — if you use the software or a derivative in a commercial product or service, you must "prominently display 'IQuest-Q1' on the user interface." Drop the commercial-attribution clause and it is plain MIT: use, modify, distribute, sell. For most research and internal use it behaves like MIT; the string only matters if IQuest-Q1 is powering something a customer sees.
What this article could verify, it did: the parameter count (reconstructed to within a million of the shard metadata), the 88-layer 3-to-1 stack and its 25 cache-growing layers (from the config, matching vLLM's allocator count), the 4-KiB-per-token KV footprint and the 3.45× saving at 512K (arithmetic on the config), the sink and partial-RoPE settings, the MTP submodel and its EAGLE serving flags, and the absence of a technical report. What it cannot: whether the benchmark numbers hold up under an independent harness, and the exact serving-kernel behavior, which lives in a fork I did not run. When the report appears — if the # ever resolves — those are the two things to check it against.