# RWKV-7 G1k: what fits in a 16-million-number mind

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/rwkv-7-g1k
> date: 2026-10-06
> tags: linear-attention, state-space-models, long-context, kv-cache, architecture, open-source, explainer

On 1 October 2026 BlinkDL [announced RWKV-7 G1k](https://x.com/BlinkDL_AI/status/2105677607568388331): "100% RNN",
"better every month", with a demo at [rwkv.com](https://rwkv.com) and weights at
[BlinkDL/rwkv7-g1](https://huggingface.co/BlinkDL/rwkv7-g1). Two days later he
[quoted his own post](https://x.com/BlinkDL_AI/status/2106418516513661189) with a stronger claim:

> Pure RNN is the safest SI approach too. All reasoning happens within its constant-size state (61x64x4096 = 16M
> numbers for RWKV-7 13B), without any growing KV cache. The tiny state is its whole mind (inner world model)
> suitable for all investigations

Three things are packed into that. An architecture: a recurrent layer whose memory is a fixed matrix per head.
An arithmetic claim: that memory is 16 million numbers for the 13B model. And an opinion: a mind that small is
the safe kind. I want to take them in that order, because the third only makes sense once the first two are
pinned down, and the second is the only one I can check exactly.

Sources: the RWKV-7 "Goose" paper ([arXiv 2503.14456](https://arxiv.org/abs/2503.14456)), the
[RWKV-LM](https://github.com/BlinkDL/RWKV-LM) repository at commit `20be0f8`, the
[Albatross](https://github.com/BlinkDL/Albatross) inference repository at `545dd18`, and the checkpoints
themselves. I did not run the models. Every number below is labelled **measured** (I read it from a file),
**reported** (someone else's figure) or **reasoned** (my arithmetic on the other two).

<ModelCard repo="BlinkDL/rwkv7-g1" />

## What shipped

The Hugging Face repository holds four G1k checkpoints, all dated `20260930` and all named `ctx25600`
(**measured**, from the file list):

| file | size on disk | layers | width | heads | parameters |
|---|---|---|---|---|---|
| `rwkv7-g1k-1.5b-…pth` | 3.06 GB | 24 | 2,048 | 32 | 1,527,668,736 |
| `rwkv7-g1k-2.9b-…pth` | 5.90 GB | 32 | 2,560 | 40 | 2,948,065,280 |
| `rwkv7-g1k-7.2b-…pth` | 14.40 GB | 32 | 4,096 | 64 | 7,199,932,416 |
| `rwkv7-g1k-13.3b-…pth` | 26.54 GB | 61 | 4,096 | 64 | 13,270,298,624 |

I did not download 50 GB to get those shapes. A `.pth` file is a zip archive whose `data.pkl` entry lists every
tensor's name and shape, so three HTTP range requests per file (the zip's central directory, the local header,
then the pickle) recover the full table. I unpickled it with a loader that refuses every class except the tensor
rebuild function, so nothing in the file executes. Layer count is the highest `blocks.N` index plus one, width is
the embedding's second dimension, and the head count is the first dimension of `att.r_k`, which is
`[heads, 64]`. Parameters are the sum over all tensors (**measured**).

What "G1k" means is also readable from the repository, once you look at its history rather than its README. The
G1 family ("GooseOne" in the model card) is one set of base models trained continuously and released as monthly
snapshots, and the letter is the snapshot. The commit log shows `g1h` on 10 July, `g1i` on 5 August, `g1j` on 31
August and `g1k` on 30 September, each upload followed by the deletion of an older letter (**measured**). The G1k
files are byte-for-byte the same size as the G1j files at every scale, so the architecture did not change; the
weights kept training (**reasoned**). The context length in the filename did change: `ctx8192` for G1g,
`ctx10240` for G1h, `ctx16384` for G1i and G1j, `ctx25600` for G1k (**measured**). I read that as the training
sequence length, which is how RWKV-LM names its checkpoints; the card does not spell it out.

The model card calls these "BASE models (pretrained with web/code/synthetic + instruction/chat/reasoning data)".
In a reply under the post, BlinkDL put the 13B's training so far at "9.8T tokens", growing by "0.83T each month
due to limited compute" (**reported**). That last clause matters for reading "better every month": the
improvement is mostly more tokens through the same network.

## The architecture in one figure

<Figure
  src="https://ai.thesatyajit.com/articles/rwkv-7-g1k/fig1.png"
  alt="Left: the RWKV-7 stack, an embedding and layer norm at the bottom, then L repeated blocks each containing a layer norm, a Time Mix module, a residual add, another layer norm, a ReLU-squared MLP and a residual add, then a final layer norm and head. Right: the inside of Time Mix. Token Shift feeds six streams labelled G, R, W, K, V, A into Weight Prepare; W, K, V and A go into the WKV7 Kernel, which takes the previous state wkv t-1 from the left and emits the new state wkv t to the right; G and R go to a Readout box."
  caption="RWKV-7's layer: a time-mix module that reads and writes a recurrent state, followed by a squared-ReLU MLP. The only things that cross from one token to the next are wkv and the token-shift inputs (RWKV-7 paper, Figure 1)."
/>

Each layer has two halves. The MLP is a plain two-matrix feed-forward with a squared ReLU and a hidden size four
times the width. The time-mix half replaces attention. It first does a **token shift**: every projection's input
is a learned per-channel blend of the current token's hidden vector $x_t$ and the previous one's $x_{t-1}$. That
costs one stored vector per layer. Then six projections come out: receptance $r$ (the query), decay $w$, key
$k$, value $v$, in-context learning rate $a$ and an output gate $g$. The **WKV kernel** uses $w, k, v, a$ to
update the state and $r$ to read it.

The interesting part is entirely in that kernel.

## The state update from first principles

Start with linear attention. Drop the softmax and attention becomes a running sum of outer products: each token
writes $v_t k_t^\top$ into a $d \times d$ matrix $S$, and a query reads $S q$. If the keys were orthonormal,
reading with key $k_i$ would return exactly $v_i$. That is a memory of fixed size, which is the point, and it
has two problems. It never forgets, so the sum keeps growing. And when you write a new value at a key that is
already in use, the old value stays: the read returns the sum of both.

Decay fixes the first problem bluntly. Multiply $S$ by a factor $w \in (0, 1)$ every step and old content fades.
RWKV-6, Mamba-2, GLA and RetNet all do some version of this. The
[KDA half-life article](/articles/kda-half-life) works out how fast a per-channel decay forgets.

The **delta rule** fixes the second problem. Before writing $v_t$ at key $k_t$, read what is already there,
$S k_t$, and subtract it, scaled by a learning rate $\beta$:

$$
S_t = S_{t-1} - \beta\,(S_{t-1} k_t)\,k_t^\top + \beta\, v_t k_t^\top = S_{t-1}\left(I - \beta\, k_t k_t^\top\right) + \beta\, v_t k_t^\top
$$

With $\beta = 1$ and a unit key, the read at $k_t$ afterwards is exactly $v_t$: the old value was erased and the
new one written. This is one step of online gradient descent on $\lVert S k_t - v_t \rVert^2$, which is why the
[linear-attention roundup](/articles/linear-attention-state-roundup) calls these layers test-time regressors.
Gated DeltaNet multiplies in a scalar decay as well.

RWKV-7 generalizes three things. Written as the paper writes it, per head:

$$
\mathbf{wkv}_t = \mathbf{wkv}_{t-1}\left(\mathrm{diag}(w_t) - \hat\kappa_t^\top (a_t \odot \hat\kappa_t)\right) + v_t^\top \tilde k_t
$$

- **Decay is a vector.** $w_t$ has one entry per key channel, computed from the token, and each entry is
  restricted to $(e^{-e^{-0.5}}, 1)$, about $(0.545, 1)$, for stability.
- **The learning rate is a vector.** $a_t = \mathrm{sigmoid}(\cdot)$ has one entry per channel, so the layer can
  replace some channels at a key and keep others. The paper calls it the in-context learning rate.
- **The key you erase at is not the key you write at.** The removal key $\hat\kappa_t$ is the key scaled by a
  learned vector and L2-normalized per head; the replacement key $\tilde k_t$ is the key blended toward
  $k_t \odot a_t$.

The reference RNN code in `RWKV-v7/rwkv_v7_demo_rnn.py` is three lines once the projections are done:

```python
# RWKV-LM @ 20be0f8, RWKV-v7/rwkv_v7_demo_rnn.py, time_mixing__
vk = v.view(H,N,1) @ k.view(H,1,N)                 # write: v k~^T
ab = (-kk).view(H,N,1) @ (kk*a).view(H,1,N)        # erase: -kappa^T (a * kappa)
state = state * w.view(H,1,N) + state @ ab.float() + vk.float()
out = state.to(dtype=x.dtype) @ r.view(H,N,1)      # read with receptance
```

`state` is `[H, N, N]` with `N = 64`: one 64 by 64 matrix per head. The paper's own picture of it is a 4 by 4
toy:

<Figure
  src="https://ai.thesatyajit.com/articles/rwkv-7-g1k/fig2.png"
  alt="A 4 by 4 coloured matrix wkv t equals the previous 4 by 4 matrix wkv t-1 multiplied by the bracket of a diagonal matrix diag of w t minus a column vector kappa hat transposed times a row vector a t elementwise kappa hat, plus a column vector v t transposed times a row vector k tilde t."
  caption="One head's update: decay each key column, erase along the removal key at a per-channel rate, add the new value at the replacement key. The real state is 64 by 64 per head (RWKV-7 paper, Figure 2)."
/>

The widget below runs that update on a 4 by 4 head, with $w$ and $a$ as scalars and $\tilde k = k$ so they fit
on sliders. It writes $v_1$ at $k_1$, $v_2$ at an overlapping key $k_2$ (their dot product is 0.36), and then
overwrites $k_1$ with $v_3$.

<DeltaStep />

With $a = 0$ and $w = 1$ (plain linear attention), reading $k_1$ after the overwrite returns $v_1 + v_3$ plus
0.36 of $v_2$: an error of 1.06 against the $v_3$ you wanted. With $a = 1$ it returns exactly $v_3$. But look at
$k_2$: erasing along $k_1$ also removed the part of $k_2$'s slot that overlapped it, and the read at $k_2$ is off
by 0.49 (linear attention is off by 0.51 there, for a different reason). A fixed matrix can only keep apart as
many things as its keys can keep apart. A 64-wide head has 64 orthogonal directions; everything beyond that
shares space (**reasoned**, from the toy).

Two properties of the transition matrix $G_t = \mathrm{diag}(w_t) - \hat\kappa_t^\top(a_t \odot \hat\kappa_t)$
are the paper's main theoretical claims (**reported**, proved in its appendix). Its eigenvalues stay in
$[-1, 1]$, so the state cannot blow up. And because $G_t$ is not diagonal and depends on the input, a stack of
these layers can track state in ways a diagonal-decay RNN or a transformer cannot under standard complexity
assumptions: the paper proves that a 4-layer RWKV-7 can recognize any regular language, and that one layer
solves the $S_5$ permutation-tracking problem. That is the strongest technical sense in which "reasoning happens
in the state" is more than a slogan: the state is not a passive buffer, it is the register of a finite automaton
the network can learn to run.

Asked under the post about Kimi Delta Attention, BlinkDL replied: "It's a weaker (and slightly faster on GPU)
form of the general DPLR RWKV-7 design… GDN => KDA => RWKV-7". DPLR is diagonal plus low rank, which describes
all three transitions: Gated DeltaNet uses a scalar decay and one key, KDA a per-channel decay and one key, RWKV-7
a per-channel decay, a per-channel learning rate and separate erase and write keys (**reasoned**, from the
three update rules). "Weaker" in the sense of strictly fewer degrees of freedom per step is fair; whether that
shows up in quality at equal compute is an empirical question nobody in that thread answered. The
[liquid time-constant article](/articles/ltc-gated-delta) traces the same recurrence back through another
literature.

## Counting the 16 million

The post's arithmetic is $61 \times 64 \times 4096$. From the checkpoint: 61 layers, 64 heads, and each head is
a $64 \times 64$ matrix, which is 4,096 numbers. So

$$
61 \times 64 \times 64 \times 64 = 15{,}990{,}784
$$

That is 15,990,784 WKV numbers (**measured** shapes, **reasoned** product). The 4,096 in the post is the head's matrix, not the
model width, which happens to also be 4,096 on this model. "16M" is right.

It is not quite the whole state. Token shift keeps the previous token's hidden vector twice per layer, once for
time-mix and once for the MLP: $2 \times 61 \times 4096$, which is 499,712 more numbers. The full recurrent state of the
13.3B model is **16,490,496 numbers** (**reasoned**). The paper's table of released models lists state size as
"WKV + Shift" in exactly this form, and its rows for the 1.5B (3,145,728 + 98,304) and 2.9B (5,242,880 +
163,840) match what I get from the G1k checkpoints at those sizes (**measured** against **reported**).

How many bytes that is depends on the runtime. RWKV-LM's reference RNN allocates the WKV state in fp32 and the
shift vectors in the model dtype; Albatross, BlinkDL's fast inference engine, keeps all of it in fp16
(**measured**, from the `torch.zeros` calls in each). For the 13.3B that is 62.0 MiB in the reference layout and
31.5 MiB in fp16 (**reasoned**).

| model | layers x heads | state numbers | reference (fp32 WKV) | fp16 |
|---|---|---|---|---|
| G1k 1.5B | 24 x 32 | 3,244,032 | 12.2 MiB | 6.2 MiB |
| G1k 2.9B | 32 x 40 | 5,406,720 | 20.3 MiB | 10.3 MiB |
| G1k 7.2B | 32 x 64 | 8,650,752 | 32.5 MiB | 16.5 MiB |
| G1k 13.3B | 61 x 64 | 16,490,496 | 62.0 MiB | 31.5 MiB |

## Against a KV cache

A transformer's equivalent is the KV cache: for every token, every layer stores a key and a value per KV head.
The cleanest comparison is the Qwen3 family, which sits in the same leaderboard and publishes its shapes. From
each model's `config.json` (**measured**), every Qwen3 dense model uses 8 KV heads of dimension 128, so the cache
per token is $2 \times \text{layers} \times 8 \times 128$ numbers: 81,920 for Qwen3-14B's 40 layers, 160 KiB
in bf16.

Divide and the crossover is early. The 13.3B's 16,490,496 state numbers equal Qwen3-14B's cache at 202 tokens;
in bytes, with the reference fp32 state, at 397 tokens (**reasoned**). At G1k's own training length of 25,600
tokens the Qwen3-14B cache is 3.9 GiB against 31.5 MiB, about 127 times larger. At $2^{20}$ tokens it is 160 GiB,
more than six times the 26.54 GB the 13.3B's weights take on disk (**reasoned**). Qwen3-14B's config lists a
40,960-token native window, so the million-token row is a hypothetical for that particular model; the per-token
rate is not. A multi-head-attention model is worse: Llama-2-13B, with 40 KV heads, stores 409,600 numbers per
token, five times Qwen3's (**measured** from its config).

<StateVsCache />

Batch size multiplies both, which is where the constant state pays off in practice. Albatross reports "10250+
token/s RWKV-7 7.2B fp16 bsz960 decoding @ RTX5090 (always const speed & vram)" (**reported**). At 16.5 MiB per
sequence in fp16, 960 sequences of state is 15.5 GiB, next to about 14.4 GB of weights: consistent with one
32 GB card (**reasoned**). The same batch of a transformer at any useful context would not fit.

Now the other side of the ledger, which the bars cannot show. A KV cache is lossless: whatever was in token 3
is still there, exactly, at token 300,000. A fixed state is lossy by construction; the widget above already shows
two overlapping keys interfering at dimension 4. The paper is candid about this. On its passkey test,
"RWKV7-World3-1.5B achieves perfect accuracy up to a context length of 19600 tokens but exhibits degradation beyond
20600 tokens", and the 2.9B holds "perfect retrieval up to 35000 tokens" before degrading (**reported**). And on
long books, the World-trained 2.9B's loss starts climbing past about 10K tokens:

<Figure
  src="https://ai.thesatyajit.com/articles/rwkv-7-g1k/fig4.png"
  alt="Line chart of average loss against token position from 1K to 32K on PG19. RWKV4-3B-World rises steeply off the top of the chart after 8K. RWKV5 and RWKV6 stay roughly flat around 2.3 and 2.25. RWKV7-2.9B-World is lowest early, around 2.13 to 2.17, then rises to about 2.25 by 30K. A RWKV7-2.9B-World3 model tuned at 128K stays lower, around 2.12 to 2.18, across the range."
  caption="PG19 loss by position. The 4K-trained RWKV-7 World model gets worse past about 10K tokens; the same model fine-tuned on long documents does not. Constant memory is not the same as unlimited recall (RWKV-7 paper, Figure 6)."
/>

The paper attributes that rise to overfitting to the 4K pretraining length and shows long-context fine-tuning
removes it. The G1 line has been pushing training length up every few months, to 25,600 for G1k, which is the
same remedy applied continuously. I have no long-context measurement for G1k itself; none was posted.

## What "better every month" measured

The release image is a screenshot of [UncheatableEval](https://huggingface.co/spaces/Jellyfish042/UncheatableEval),
a leaderboard that scores base models by how well they compress text published after their training cutoffs
(lower is better; the image does not name the unit, and compression ratio is the first of the metrics the leaderboard's code offers).

<Figure
  src="https://ai.thesatyajit.com/articles/rwkv-7-g1k/fig5.jpg"
  alt="Two leaderboard tables. Top: about two dozen base models sorted by average compression score, lower is better, with columns for GitHub code, arXiv, bioRxiv, Wikipedia, BBC news and AO3. gemma-4-31B leads at 6.001, then Qwen3.5-35B-A3B-Base 6.242, gemma-4-26B-A4B 6.259, Mistral-Small-3.1-24B-Base 6.284, then the highlighted rwkv7-g1k-13.3b at 6.288, gemma-3-27b-pt 6.302, rwkv7-g1j-13.3b 6.313, rwkv7-g1i 6.343, rwkv7-g1h 6.362. Bottom: a second table of models up to about 16B parameters, averaged over the code, arXiv and bioRxiv columns only, where rwkv7-g1k-13.3b leads at 5.373 ahead of g1j 5.401, Ministral-3-14B 5.424 and Qwen3.5-9B-Base 5.434."
  caption="The G1k release image. Top, all data sources; bottom, smaller models only, scored on code, arXiv and bioRxiv. BlinkDL's note between them reads that Qwen 3.5 is still ahead on code and arXiv (BlinkDL's G1k release post on X, image 1)."
/>

Read off the image (**reported**): the 13.3B scores 6.288 for G1k, 6.313 for G1j, 6.343 for G1i and 6.362 for
G1h, so the monthly steps are 0.019, 0.030 and 0.025 (**reasoned**). G1k lands between Mistral-Small-3.1-24B
(6.284) and gemma-3-27b (6.302), behind gemma-4-31B (6.001) and Qwen3.5-35B-A3B (6.242), all of them larger. On
the lower table, which keeps only models up to about 16B and only the code, arXiv and bioRxiv columns, it leads at
5.373, ahead of Ministral-3-14B (5.424) and Qwen3.5-9B (5.434). BlinkDL's own caption
concedes the gap: "Qwen 3.5 is still ahead in code + arxiv".

Three cautions on reading this. It is a compression benchmark of base models, so it says nothing directly about
chat, tool use or reasoning; the model card links a GPQA evaluation script but the release posted no such
number. The lower table is a subset of models and columns chosen by the publisher; it is not wrong, but it is the
flattering cut. And "better every month" is literally true on this metric for four consecutive releases, at a cost of
about 0.83T tokens a month.

For the architecture itself, the paper's headline evidence is older and on smaller models: trained on far fewer
tokens than its peers, RWKV7-World3 traces a better accuracy-per-FLOP curve on multilingual benchmarks than
Qwen2.5 and SmolLM2 (**reported**).

<Figure
  src="https://ai.thesatyajit.com/articles/rwkv-7-g1k/fig3.png"
  alt="Scatter with lines: average multilingual benchmark accuracy against training FLOPs on a log scale. RWKV7-World3 models at 0.19B, 0.4B, 1.5B and 2.9B form a red line from about 47.5 to 61 percent, to the left of and above the blue Qwen2.5 line from 0.5B to 7B and the green SmolLM2 line from 135M to 1.7B."
  caption="Multilingual accuracy against training compute for RWKV7-World3, Qwen2.5 and SmolLM2. The x axis is training compute, so a model trained on fewer tokens moves left (RWKV-7 paper, Figure 3a)."
/>

## The safety claim, calmly

"The safest SI approach" is an opinion, and the post does not argue it beyond the size of the state. It is worth
separating what in it is true, what is overstated, and what could be tested.

**What is true.** Everything an RWKV-7 model carries from one token to the next is a fixed-size object you can
copy, save, diff, restore and edit. For the 13.3B, 16,490,496 numbers; in fp16 that is at most about 33 million
bytes of information per sequence, no matter how long the sequence runs (**reasoned**). That is a hard ceiling a
transformer does not have, and it is an operational convenience as well: you can checkpoint a conversation as
31.5 MiB, fork it, or roll it back. The ecosystem already edits states directly; [RWKV-PEFT](https://github.com/Joluck/RWKV-PEFT),
which the model card links for training, lists "State Tuning" next to LoRA, which learns an initial state
instead of new weights.

**What is overstated.** "All reasoning happens within its constant-size state" is true of what persists across
tokens, not of computation. The reasoning on any single token happens in 13.3 billion weights, and the state is
only the thread between steps. A transformer's KV cache is also a complete, finite, inspectable record of
everything the model carries forward; it is bigger, and it grows, but it is not less transparent. And a model that
writes a chain of thought carries part of its "mind" in the emitted tokens, which it reads back as input, for an
RNN as for a transformer. 16 million numbers is small next to a 160 GiB cache, but it is not small next to what
interpretability tools currently explain; nobody can read a 16M-dimensional vector as a world model yet.

**What is testable.** Each of these would turn the opinion into evidence, and none needs anything beyond the
released weights:

1. **Probes.** Train linear probes on the WKV state to recover facts about the context: which entities were
   mentioned, a board position, a variable's value in code. The paper's appendix already trains small RWKV-7
   models on Reversi and reports they learn board state tracking before evaluation; a probe on the state would show
   where that board lives.
2. **Interventions.** Patch one head's 64 by 64 matrix from another context and see whether the downstream
   behaviour changes the way the probe predicts. If editing the state reliably edits beliefs, "the state is its
   whole mind" earns its keep.
3. **Capacity.** Measure how recall degrades with the number of distinct facts held, against the 64-direction
   ceiling per head the toy above illustrates. That bounds what any long-horizon plan can keep without writing it
   down.

If those come back clean, a bounded, editable state is a real advantage for oversight over a cache that grows
without limit. If they come back like most interpretability results, it is an advantage of degree.

## What I would take away

The 16M figure is right, and checkable from the checkpoint without running it: 61 layers, 64 heads, 64 by 64
per head, plus half a million numbers of token shift. The update that fills that state is a delta rule made
per-channel in both forgetting and learning, with separate keys for erasing and writing, and the paper's
expressivity results are about exactly that structure. Against a same-size transformer the state is the size of a
few hundred tokens of KV cache, which is why RWKV serves hundreds of sequences on one consumer GPU. The price is
that the memory is lossy, and the long-context curves show it. The safety argument is an interesting hypothesis
about oversight, not a result; the experiments that would settle it are cheap, and the weights are Apache-2.0.
