~/satyajit

RWKV-7 G1k: what fits in a 16-million-number mind

mdjsonmcp

2026-10-06 · 19 min · linear-attention · state-space-models · long-context · kv-cache · architecture · open-source · explainer

On 1 October 2026 BlinkDL announced RWKV-7 G1k: "100% RNN", "better every month", with a demo at rwkv.com and weights at BlinkDL/rwkv7-g1. Two days later he quoted his own post with a stronger claim:

Pure RNN is the safest SI approach too. All reasoning happens within its constant-size state (61x64x4096 = 16M numbers for RWKV-7 13B), without any growing KV cache. The tiny state is its whole mind (inner world model) suitable for all investigations

Three things are packed into that. An architecture: a recurrent layer whose memory is a fixed matrix per head. An arithmetic claim: that memory is 16 million numbers for the 13B model. And an opinion: a mind that small is the safe kind. I want to take them in that order, because the third only makes sense once the first two are pinned down, and the second is the only one I can check exactly.

Sources: the RWKV-7 "Goose" paper (arXiv 2503.14456), the RWKV-LM repository at commit 20be0f8, the Albatross inference repository at 545dd18, and the checkpoints themselves. I did not run the models. Every number below is labelled measured (I read it from a file), reported (someone else's figure) or reasoned (my arithmetic on the other two).

BlinkDL/rwkv7-g1@cd67fb9 · snapshot 2026-10-06
repo size
671.35 GB
task
text-generation
library
rwkv
license
apache-2.0
largest file
26.54 GB
files
14
downloads
25.8K
likes
231
languages
en, zh, fr, es, de, pt
pytorchtext-generationcausal-lmrwkv

repo last modified 2026-09-30

What shipped

The Hugging Face repository holds four G1k checkpoints, all dated 20260930 and all named ctx25600 (measured, from the file list):

filesize on disklayerswidthheadsparameters
rwkv7-g1k-1.5b-…pth3.06 GB242,048321,527,668,736
rwkv7-g1k-2.9b-…pth5.90 GB322,560402,948,065,280
rwkv7-g1k-7.2b-…pth14.40 GB324,096647,199,932,416
rwkv7-g1k-13.3b-…pth26.54 GB614,0966413,270,298,624

I did not download 50 GB to get those shapes. A .pth file is a zip archive whose data.pkl entry lists every tensor's name and shape, so three HTTP range requests per file (the zip's central directory, the local header, then the pickle) recover the full table. I unpickled it with a loader that refuses every class except the tensor rebuild function, so nothing in the file executes. Layer count is the highest blocks.N index plus one, width is the embedding's second dimension, and the head count is the first dimension of att.r_k, which is [heads, 64]. Parameters are the sum over all tensors (measured).

What "G1k" means is also readable from the repository, once you look at its history rather than its README. The G1 family ("GooseOne" in the model card) is one set of base models trained continuously and released as monthly snapshots, and the letter is the snapshot. The commit log shows g1h on 10 July, g1i on 5 August, g1j on 31 August and g1k on 30 September, each upload followed by the deletion of an older letter (measured). The G1k files are byte-for-byte the same size as the G1j files at every scale, so the architecture did not change; the weights kept training (reasoned). The context length in the filename did change: ctx8192 for G1g, ctx10240 for G1h, ctx16384 for G1i and G1j, ctx25600 for G1k (measured). I read that as the training sequence length, which is how RWKV-LM names its checkpoints; the card does not spell it out.

The model card calls these "BASE models (pretrained with web/code/synthetic + instruction/chat/reasoning data)". In a reply under the post, BlinkDL put the 13B's training so far at "9.8T tokens", growing by "0.83T each month due to limited compute" (reported). That last clause matters for reading "better every month": the improvement is mostly more tokens through the same network.

The architecture in one figure

Left: the RWKV-7 stack, an embedding and layer norm at the bottom, then L repeated blocks each containing a layer norm, a Time Mix module, a residual add, another layer norm, a ReLU-squared MLP and a residual add, then a final layer norm and head. Right: the inside of Time Mix. Token Shift feeds six streams labelled G, R, W, K, V, A into Weight Prepare; W, K, V and A go into the WKV7 Kernel, which takes the previous state wkv t-1 from the left and emits the new state wkv t to the right; G and R go to a Readout box.
RWKV-7's layer: a time-mix module that reads and writes a recurrent state, followed by a squared-ReLU MLP. The only things that cross from one token to the next are wkv and the token-shift inputs (RWKV-7 paper, Figure 1).

Each layer has two halves. The MLP is a plain two-matrix feed-forward with a squared ReLU and a hidden size four times the width. The time-mix half replaces attention. It first does a token shift: every projection's input is a learned per-channel blend of the current token's hidden vector xtx_t and the previous one's xt−1x_{t-1}. That costs one stored vector per layer. Then six projections come out: receptance rr (the query), decay ww, key kk, value vv, in-context learning rate aa and an output gate gg. The WKV kernel uses w,k,v,aw, k, v, a to update the state and rr to read it.

The interesting part is entirely in that kernel.

The state update from first principles

Start with linear attention. Drop the softmax and attention becomes a running sum of outer products: each token writes vtkt⊤v_t k_t^\top into a d×dd \times d matrix SS, and a query reads SqS q. If the keys were orthonormal, reading with key kik_i would return exactly viv_i. That is a memory of fixed size, which is the point, and it has two problems. It never forgets, so the sum keeps growing. And when you write a new value at a key that is already in use, the old value stays: the read returns the sum of both.

Decay fixes the first problem bluntly. Multiply SS by a factor w∈(0,1)w \in (0, 1) every step and old content fades. RWKV-6, Mamba-2, GLA and RetNet all do some version of this. The KDA half-life article works out how fast a per-channel decay forgets.

The delta rule fixes the second problem. Before writing vtv_t at key ktk_t, read what is already there, SktS k_t, and subtract it, scaled by a learning rate β\beta:

St=St−1−β (St−1kt) kt⊤+β vtkt⊤=St−1(I−β ktkt⊤)+β vtkt⊤S_t = S_{t-1} - \beta\,(S_{t-1} k_t)\,k_t^\top + \beta\, v_t k_t^\top = S_{t-1}\left(I - \beta\, k_t k_t^\top\right) + \beta\, v_t k_t^\top

With β=1\beta = 1 and a unit key, the read at ktk_t afterwards is exactly vtv_t: the old value was erased and the new one written. This is one step of online gradient descent on ∥Skt−vt∥2\lVert S k_t - v_t \rVert^2, which is why the linear-attention roundup calls these layers test-time regressors. Gated DeltaNet multiplies in a scalar decay as well.

RWKV-7 generalizes three things. Written as the paper writes it, per head:

wkvt=wkvt−1(diag(wt)−κ^t⊤(at⊙κ^t))+vt⊤k~t\mathbf{wkv}_t = \mathbf{wkv}_{t-1}\left(\mathrm{diag}(w_t) - \hat\kappa_t^\top (a_t \odot \hat\kappa_t)\right) + v_t^\top \tilde k_t

The reference RNN code in RWKV-v7/rwkv_v7_demo_rnn.py is three lines once the projections are done:

# RWKV-LM @ 20be0f8, RWKV-v7/rwkv_v7_demo_rnn.py, time_mixing__
vk = v.view(H,N,1) @ k.view(H,1,N)                 # write: v k~^T
ab = (-kk).view(H,N,1) @ (kk*a).view(H,1,N)        # erase: -kappa^T (a * kappa)
state = state * w.view(H,1,N) + state @ ab.float() + vk.float()
out = state.to(dtype=x.dtype) @ r.view(H,N,1)      # read with receptance

state is [H, N, N] with N = 64: one 64 by 64 matrix per head. The paper's own picture of it is a 4 by 4 toy:

A 4 by 4 coloured matrix wkv t equals the previous 4 by 4 matrix wkv t-1 multiplied by the bracket of a diagonal matrix diag of w t minus a column vector kappa hat transposed times a row vector a t elementwise kappa hat, plus a column vector v t transposed times a row vector k tilde t.
One head's update: decay each key column, erase along the removal key at a per-channel rate, add the new value at the replacement key. The real state is 64 by 64 per head (RWKV-7 paper, Figure 2).

The widget below runs that update on a 4 by 4 head, with ww and aa as scalars and k~=k\tilde k = k so they fit on sliders. It writes v1v_1 at k1k_1, v2v_2 at an overlapping key k2k_2 (their dot product is 0.36), and then overwrites k1k_1 with v3v_3.

one head, 4×4 instead of 64×64
state S after 3 writes
0.10
-0.14
-0.29
0.00
-0.29
0.38
0.80
0.00
0.80
0.60
0.00
0.00
0.00
0.00
0.00
0.00
rows: value dims · columns: key dims
  1. 1. write v₁ at key k₁
  2. 2. write v₂ at key k₂ (k₁·k₂ = 0.36, they overlap)
  3. 3. overwrite: write v₃ at key k₁ again
ready = S rwanterror
r = k₁[0.00, 0.00, 1.00, 0.00]v₃0.00
r = k₂[-0.31, 0.87, 0.36, 0.00]v₂0.49
writes3
learning rate a1.00
decay w1.00

With a = 0 the overwrite stacks on top: reading k₁ returns v₁ + v₃ plus 0.36 of v₂. With a = 1 it returns exactly v₃, but the erase at k₁ also takes 0.36 of the overlap out of k₂’s slot. A fixed state stores cleanly only what its keys can keep apart.

With a=0a = 0 and w=1w = 1 (plain linear attention), reading k1k_1 after the overwrite returns v1+v3v_1 + v_3 plus 0.36 of v2v_2: an error of 1.06 against the v3v_3 you wanted. With a=1a = 1 it returns exactly v3v_3. But look at k2k_2: erasing along k1k_1 also removed the part of k2k_2's slot that overlapped it, and the read at k2k_2 is off by 0.49 (linear attention is off by 0.51 there, for a different reason). A fixed matrix can only keep apart as many things as its keys can keep apart. A 64-wide head has 64 orthogonal directions; everything beyond that shares space (reasoned, from the toy).

Two properties of the transition matrix Gt=diag(wt)−κ^t⊤(at⊙κ^t)G_t = \mathrm{diag}(w_t) - \hat\kappa_t^\top(a_t \odot \hat\kappa_t) are the paper's main theoretical claims (reported, proved in its appendix). Its eigenvalues stay in [−1,1][-1, 1], so the state cannot blow up. And because GtG_t is not diagonal and depends on the input, a stack of these layers can track state in ways a diagonal-decay RNN or a transformer cannot under standard complexity assumptions: the paper proves that a 4-layer RWKV-7 can recognize any regular language, and that one layer solves the S5S_5 permutation-tracking problem. That is the strongest technical sense in which "reasoning happens in the state" is more than a slogan: the state is not a passive buffer, it is the register of a finite automaton the network can learn to run.

Asked under the post about Kimi Delta Attention, BlinkDL replied: "It's a weaker (and slightly faster on GPU) form of the general DPLR RWKV-7 design… GDN => KDA => RWKV-7". DPLR is diagonal plus low rank, which describes all three transitions: Gated DeltaNet uses a scalar decay and one key, KDA a per-channel decay and one key, RWKV-7 a per-channel decay, a per-channel learning rate and separate erase and write keys (reasoned, from the three update rules). "Weaker" in the sense of strictly fewer degrees of freedom per step is fair; whether that shows up in quality at equal compute is an empirical question nobody in that thread answered. The liquid time-constant article traces the same recurrence back through another literature.

Counting the 16 million

The post's arithmetic is 61×64×409661 \times 64 \times 4096. From the checkpoint: 61 layers, 64 heads, and each head is a 64×6464 \times 64 matrix, which is 4,096 numbers. So

61×64×64×64=15,990,78461 \times 64 \times 64 \times 64 = 15{,}990{,}784

That is 15,990,784 WKV numbers (measured shapes, reasoned product). The 4,096 in the post is the head's matrix, not the model width, which happens to also be 4,096 on this model. "16M" is right.

It is not quite the whole state. Token shift keeps the previous token's hidden vector twice per layer, once for time-mix and once for the MLP: 2×61×40962 \times 61 \times 4096, which is 499,712 more numbers. The full recurrent state of the 13.3B model is 16,490,496 numbers (reasoned). The paper's table of released models lists state size as "WKV + Shift" in exactly this form, and its rows for the 1.5B (3,145,728 + 98,304) and 2.9B (5,242,880 + 163,840) match what I get from the G1k checkpoints at those sizes (measured against reported).

How many bytes that is depends on the runtime. RWKV-LM's reference RNN allocates the WKV state in fp32 and the shift vectors in the model dtype; Albatross, BlinkDL's fast inference engine, keeps all of it in fp16 (measured, from the torch.zeros calls in each). For the 13.3B that is 62.0 MiB in the reference layout and 31.5 MiB in fp16 (reasoned).

modellayers x headsstate numbersreference (fp32 WKV)fp16
G1k 1.5B24 x 323,244,03212.2 MiB6.2 MiB
G1k 2.9B32 x 405,406,72020.3 MiB10.3 MiB
G1k 7.2B32 x 648,650,75232.5 MiB16.5 MiB
G1k 13.3B61 x 6416,490,49662.0 MiB31.5 MiB

Against a KV cache

A transformer's equivalent is the KV cache: for every token, every layer stores a key and a value per KV head. The cleanest comparison is the Qwen3 family, which sits in the same leaderboard and publishes its shapes. From each model's config.json (measured), every Qwen3 dense model uses 8 KV heads of dimension 128, so the cache per token is 2×layers×8×1282 \times \text{layers} \times 8 \times 128 numbers: 81,920 for Qwen3-14B's 40 layers, 160 KiB in bf16.

Divide and the crossover is early. The 13.3B's 16,490,496 state numbers equal Qwen3-14B's cache at 202 tokens; in bytes, with the reference fp32 state, at 397 tokens (reasoned). At G1k's own training length of 25,600 tokens the Qwen3-14B cache is 3.9 GiB against 31.5 MiB, about 127 times larger. At 2202^{20} tokens it is 160 GiB, more than six times the 26.54 GB the 13.3B's weights take on disk (reasoned). Qwen3-14B's config lists a 40,960-token native window, so the million-token row is a hypothetical for that particular model; the per-token rate is not. A multi-head-attention model is worse: Llama-2-13B, with 40 KV heads, stores 409,600 numbers per token, five times Qwen3's (measured from its config).

memory carried between tokens · log scale
RWKV-7 13.3B state62 MiB · same at every length
Qwen3-14B KV cache5 GiB · grows with every token
RWKV-7 13.3B weights, bf1624.7 GiB · for scale, shared by the batch
context length32K tokens
concurrent sequences1
state numbers
61·64·64·64 + 2·61·4096 = 16,490,496
KV numbers per token
2·40·8·128 = 81,920
equal in numbers at
202 tokens
equal in bytes at
397 tokens
cache ÷ state at this length
82.6x

The RWKV bar does not move with the context slider; that is the whole claim. The KV bar crosses it within a few hundred tokens for every size, and by a million tokens a 14B-class cache is larger than the 13.3B model’s own weights. What the bars do not show is what each memory can hold: the cache keeps every key and value exactly, the state keeps whatever the update rule decided was worth keeping.

Batch size multiplies both, which is where the constant state pays off in practice. Albatross reports "10250+ token/s RWKV-7 7.2B fp16 bsz960 decoding @ RTX5090 (always const speed & vram)" (reported). At 16.5 MiB per sequence in fp16, 960 sequences of state is 15.5 GiB, next to about 14.4 GB of weights: consistent with one 32 GB card (reasoned). The same batch of a transformer at any useful context would not fit.

Now the other side of the ledger, which the bars cannot show. A KV cache is lossless: whatever was in token 3 is still there, exactly, at token 300,000. A fixed state is lossy by construction; the widget above already shows two overlapping keys interfering at dimension 4. The paper is candid about this. On its passkey test, "RWKV7-World3-1.5B achieves perfect accuracy up to a context length of 19600 tokens but exhibits degradation beyond 20600 tokens", and the 2.9B holds "perfect retrieval up to 35000 tokens" before degrading (reported). And on long books, the World-trained 2.9B's loss starts climbing past about 10K tokens:

Line chart of average loss against token position from 1K to 32K on PG19. RWKV4-3B-World rises steeply off the top of the chart after 8K. RWKV5 and RWKV6 stay roughly flat around 2.3 and 2.25. RWKV7-2.9B-World is lowest early, around 2.13 to 2.17, then rises to about 2.25 by 30K. A RWKV7-2.9B-World3 model tuned at 128K stays lower, around 2.12 to 2.18, across the range.
PG19 loss by position. The 4K-trained RWKV-7 World model gets worse past about 10K tokens; the same model fine-tuned on long documents does not. Constant memory is not the same as unlimited recall (RWKV-7 paper, Figure 6).

The paper attributes that rise to overfitting to the 4K pretraining length and shows long-context fine-tuning removes it. The G1 line has been pushing training length up every few months, to 25,600 for G1k, which is the same remedy applied continuously. I have no long-context measurement for G1k itself; none was posted.

What "better every month" measured

The release image is a screenshot of UncheatableEval, a leaderboard that scores base models by how well they compress text published after their training cutoffs (lower is better; the image does not name the unit, and compression ratio is the first of the metrics the leaderboard's code offers).

Two leaderboard tables. Top: about two dozen base models sorted by average compression score, lower is better, with columns for GitHub code, arXiv, bioRxiv, Wikipedia, BBC news and AO3. gemma-4-31B leads at 6.001, then Qwen3.5-35B-A3B-Base 6.242, gemma-4-26B-A4B 6.259, Mistral-Small-3.1-24B-Base 6.284, then the highlighted rwkv7-g1k-13.3b at 6.288, gemma-3-27b-pt 6.302, rwkv7-g1j-13.3b 6.313, rwkv7-g1i 6.343, rwkv7-g1h 6.362. Bottom: a second table of models up to about 16B parameters, averaged over the code, arXiv and bioRxiv columns only, where rwkv7-g1k-13.3b leads at 5.373 ahead of g1j 5.401, Ministral-3-14B 5.424 and Qwen3.5-9B-Base 5.434.
The G1k release image. Top, all data sources; bottom, smaller models only, scored on code, arXiv and bioRxiv. BlinkDL's note between them reads that Qwen 3.5 is still ahead on code and arXiv (BlinkDL's G1k release post on X, image 1).

Read off the image (reported): the 13.3B scores 6.288 for G1k, 6.313 for G1j, 6.343 for G1i and 6.362 for G1h, so the monthly steps are 0.019, 0.030 and 0.025 (reasoned). G1k lands between Mistral-Small-3.1-24B (6.284) and gemma-3-27b (6.302), behind gemma-4-31B (6.001) and Qwen3.5-35B-A3B (6.242), all of them larger. On the lower table, which keeps only models up to about 16B and only the code, arXiv and bioRxiv columns, it leads at 5.373, ahead of Ministral-3-14B (5.424) and Qwen3.5-9B (5.434). BlinkDL's own caption concedes the gap: "Qwen 3.5 is still ahead in code + arxiv".

Three cautions on reading this. It is a compression benchmark of base models, so it says nothing directly about chat, tool use or reasoning; the model card links a GPQA evaluation script but the release posted no such number. The lower table is a subset of models and columns chosen by the publisher; it is not wrong, but it is the flattering cut. And "better every month" is literally true on this metric for four consecutive releases, at a cost of about 0.83T tokens a month.

For the architecture itself, the paper's headline evidence is older and on smaller models: trained on far fewer tokens than its peers, RWKV7-World3 traces a better accuracy-per-FLOP curve on multilingual benchmarks than Qwen2.5 and SmolLM2 (reported).

Scatter with lines: average multilingual benchmark accuracy against training FLOPs on a log scale. RWKV7-World3 models at 0.19B, 0.4B, 1.5B and 2.9B form a red line from about 47.5 to 61 percent, to the left of and above the blue Qwen2.5 line from 0.5B to 7B and the green SmolLM2 line from 135M to 1.7B.
Multilingual accuracy against training compute for RWKV7-World3, Qwen2.5 and SmolLM2. The x axis is training compute, so a model trained on fewer tokens moves left (RWKV-7 paper, Figure 3a).

The safety claim, calmly

"The safest SI approach" is an opinion, and the post does not argue it beyond the size of the state. It is worth separating what in it is true, what is overstated, and what could be tested.

What is true. Everything an RWKV-7 model carries from one token to the next is a fixed-size object you can copy, save, diff, restore and edit. For the 13.3B, 16,490,496 numbers; in fp16 that is at most about 33 million bytes of information per sequence, no matter how long the sequence runs (reasoned). That is a hard ceiling a transformer does not have, and it is an operational convenience as well: you can checkpoint a conversation as 31.5 MiB, fork it, or roll it back. The ecosystem already edits states directly; RWKV-PEFT, which the model card links for training, lists "State Tuning" next to LoRA, which learns an initial state instead of new weights.

What is overstated. "All reasoning happens within its constant-size state" is true of what persists across tokens, not of computation. The reasoning on any single token happens in 13.3 billion weights, and the state is only the thread between steps. A transformer's KV cache is also a complete, finite, inspectable record of everything the model carries forward; it is bigger, and it grows, but it is not less transparent. And a model that writes a chain of thought carries part of its "mind" in the emitted tokens, which it reads back as input, for an RNN as for a transformer. 16 million numbers is small next to a 160 GiB cache, but it is not small next to what interpretability tools currently explain; nobody can read a 16M-dimensional vector as a world model yet.

What is testable. Each of these would turn the opinion into evidence, and none needs anything beyond the released weights:

  1. Probes. Train linear probes on the WKV state to recover facts about the context: which entities were mentioned, a board position, a variable's value in code. The paper's appendix already trains small RWKV-7 models on Reversi and reports they learn board state tracking before evaluation; a probe on the state would show where that board lives.
  2. Interventions. Patch one head's 64 by 64 matrix from another context and see whether the downstream behaviour changes the way the probe predicts. If editing the state reliably edits beliefs, "the state is its whole mind" earns its keep.
  3. Capacity. Measure how recall degrades with the number of distinct facts held, against the 64-direction ceiling per head the toy above illustrates. That bounds what any long-horizon plan can keep without writing it down.

If those come back clean, a bounded, editable state is a real advantage for oversight over a cache that grows without limit. If they come back like most interpretability results, it is an advantage of degree.

What I would take away

The 16M figure is right, and checkable from the checkpoint without running it: 61 layers, 64 heads, 64 by 64 per head, plus half a million numbers of token shift. The update that fills that state is a delta rule made per-channel in both forgetting and learning, with separate keys for erasing and writing, and the paper's expressivity results are about exactly that structure. Against a same-size transformer the state is the size of a few hundred tokens of KV cache, which is why RWKV serves hundreds of sequences on one consumer GPU. The price is that the memory is lossy, and the long-context curves show it. The safety argument is an interesting hypothesis about oversight, not a result; the experiments that would settle it are cheap, and the weights are Apache-2.0.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "RWKV-7 G1k: what fits in a 16-million-number mind", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026rwkv7g1k,
  author = {Satyajit Ghana},
  title  = {RWKV-7 G1k: what fits in a 16-million-number mind},
  url    = {https://ai.thesatyajit.com/articles/rwkv-7-g1k},
  year   = {2026}
}
share