~/satyajit

Token cues: how 'chicken' made a base model reason as well as RL

mdjsonmcp

2026-10-06 · 27 min · reinforcement-learning · reasoning · pretraining · interpretability

Why read this

Notabletop 60%

How two forced tokens match RL-Zero on Olmo, the swap test that puts RL's gain in the opening, the chicken data edit, and where the result stops.

  • Runs on a consumer GPU
  • Widely used
  • A lasting reference

Training & RLResearch paper

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
1 of 3: Spot-checks a few numbers
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
2 of 3: A widely used model, tool or lab release
Only here?
1 of 3: Some original analysis

Score 60 of 100, ranked 231 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

The post that sent me here said a base model could be made to reason with a chicken. I assumed it was a prompt-golf result, the sort where someone finds a weird string that bumps one benchmark on one model. It is a better paper than that. The chicken is the last experiment in a chain, and the chain says something uncomfortable about what reinforcement learning on reasoning has been doing.

The paper is Base Models Can Reason By Taking a Cue From Training Data by Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min and Alexei A. Efros (MIT, Berkeley, UW, Ai2; 5 October 2026). Its code is on GitHub. It came to me through Kevin Farhat's post quoting Sophie Wang's thread. The post's line was: swap "okay" for "chicken" in training data, and starting with .\n\nChicken lifts MATH-500 from 18% to 77%, matching .\n\nOkay and RL.

That sentence is accurate, with one thing it leaves out. The 18% to 77% is the full-scale rerun in the appendix (Table 17: 17.6% on the released base, 76.9% after the edited 100B-token mid-training). The main experiment in Section 3 runs a tenth of that budget, where the same edit takes the chicken cue from 2.4% to 37.2%. Both are real. They are two different training runs.

Diagram. A math prompt feeds a token tree. From the prompt, the base model's likely first tokens are '.\n' with probability 0.38, then 'Answer' 0.10, leading to a concise wrong answer of 1002 that the figure links to short Q&A training data. The other branch is '.\n\n' with 0.24, then 'Okay' with 0.57, leading to a worked solution that checks itself and answers 1251 correctly, linked to reasoning traces in the training data. A red line marks 'Chicken' with probability 2e-7, which the data edit raises. A blue line marks RL raising the probability of the Okay branch.
The whole argument on one page: the base model's first tokens pick between a short-answer habit and a reasoning-trace habit, RL makes the good branch likely, and a data edit can make an arbitrary word into the good branch (paper, Figure 1).

What two tokens can do

Start with the model's own factorisation. Given a prompt xx, the probability of an opening cc and then a continuation yy splits in two:

pθ(c,y∣x)=pθ(c∣x) pθ(y∣x,c)p_\theta(c, y \mid x) = p_\theta(c \mid x)\, p_\theta(y \mid x, c)

Normally the model samples both. The paper's trick is to stop sampling cc and write it in yourself, a prefill, so you see pθ(y∣x,c)p_\theta(y \mid x, c) on its own. No weights change. No instruction is added.

For Olmo-3-7B under Ai2's RL-Zero prompt, the opening is .\n\nOkay: a period, a blank line, and the word "Okay". That looks like nothing until you read the prompt. In the code it is cues/prompts.py:22-26, and it ends with Remember to put your answer on its own line after "Answer:" and no full stop. So the period finishes the instruction, the blank line starts a new paragraph, and "Okay" is the first word of the model's answer. The cue is the shape of how a particular kind of document begins.

The authors noticed it from correlations first. In 2,000 uncued Olmo responses to MATH-500, a paragraph break opens 50% of the correct answers and 14% of the wrong ones (Appendix C.8). Then they forced it.

Bar charts. Left: MATH-500 pass@1. Olmo-3-7B base no cue 0.42, base with '.\n\n Okay' 0.78, after RL 0.75. Qwen3-14B base no cue 0.72, base with ' Alright,' 0.87, after RL 0.87. Right: four small panels for GSM8K, AMC 23, AIME 2024 and HumanEval with the same three bars per model, the cued bar usually at or above the RL bar.
Two fixed tokens against a full RL run. Olmo-3-7B goes from 42% to 78% with the cue, against 75% for its RL-Zero model; Qwen3-14B goes from 72% to 87%, the same as its GRPO model (paper, Figure 2).

Olmo-3-7B goes from 42% to 78% on MATH-500. The RL-Zero model Ai2 released gets 75%. Qwen3-14B, with its own cue Alright,, goes from 72% to 87%, which is exactly where the authors' GRPO-trained Qwen lands. The same two tokens, chosen on MATH training problems, carry over unchanged to GSM8K, AMC 23 and AIME 2024. On HumanEval, Olmo's cue lifts code from 50% to 70%; Qwen's leaves it near its base accuracy, as RL does.

The example the paper uses is a school problem: 834 students take music, they are two-thirds of the school, how many students are there? With .\nAnswer prefilled, Olmo writes "Answer: 1002". With .\n\nOkay it writes "(2/3) * T = 834", gets 1251, and then, unprompted, "Let me double-check to make sure I didn't make a mistake." That self-check is the behaviour the DeepSeek-R1-Zero paper called an "aha moment" and attributed to RL. Here it came out of a base model handed two tokens.

A detail I like: the paragraph break on its own scores 76.6% (Appendix A.2.2). The reason is that 95% of sampled responses that start with .\n\n go on to write "Okay" anyway. The bare word Okay gets 74.9% when forced, but sits 542nd in the model's ranking of first tokens, so a search that starts from what the model would plausibly say never proposes it. The cue works because it is a door the model was already standing next to.

Finding the cue without the answer key

If you could pick the opening by trying every candidate against MATH-500 labels, this would be a much weaker result: you would have tuned two tokens on the test. The search does not do that, and the code is careful about it.

It has two stages. First, a beam search over the model's own first two tokens after the prompt, pooled over MATH training problems, keeping the 20 openings with the most probability mass. Then each of those openings is screened: 16 sampled answers per problem on 30 training problems, and the winner is the one whose answers agree with each other most, measured as the mean Shannon entropy of the answer distribution. A missing answer counts as its own answer, and any opening whose missing-answer rate is more than 10 points above no-cue is dropped, so a cue that makes the model go quiet cannot pass for a confident one. The decision is four lines in cues/agree.py:95-101:

# cues/agree.py:95-101
def decide(scores, guard=0.10):
    """Mark eligibility and return the winning slug (lowest entropy among the eligible), or None."""
    base = scores.get("none")
    for slug, s in scores.items():
        s["eligible"] = slug != "none" and (base is None or s["none_share"] <= base["none_share"] + guard)
    eligible = [k for k, s in scores.items() if s["eligible"]]
    return min(eligible, key=lambda k: (scores[k]["entropy"], k)) if eligible else None

Answer agreement as a stand-in for accuracy is borrowed from self-consistency work, and it held up here. When the authors brute-forced 500 two-token openings on MATH-500 (the 50 likeliest first tokens times their 10 likeliest successors), the best scored 78.4% and the label-free pick scored 77.3%, eighth of 501 (Appendix A.2.1). Random openings went the other way: a paragraph break followed by a random vocabulary token scored 33.3% to 47.1% across five draws, against the cue's 76.9%.

One thing in the code does not match the paper's text. Section 2.1 says the beam ranks openings averaged over 30 training problems with beam width 20. The released cues/beam.py:44-49 defaults to 100 probe problems, a beam width of 40, 12 first tokens and 4 continuations each, keeping openings that appear on at least half the problems; cues/search.py then takes the top 20 by mass and screens them on the first 30 problems. The 20 and the 30 in the paper are the screen, and the beam is wider than described. It does not change the result, but if you reimplement from the paper you will search a smaller space than they did.

The search is also not cheap. The docstring in cues/search.py:13-15 puts it at about 4 GPU-hours for Qwen3-14B and about 25 for Olmo-3-7B, whose base rollouts run long: 21 openings times 480 rollouts at a 16k-token budget. Still far below an RL run, and it needs no reward model, no gradient and no labels.

The paper also pits it against GEPA, a prompt optimiser that gets correctness feedback and a GPT-4.1-mini reflector. GEPA found a 9-token prefill for Olmo and a 64-token one for Qwen, both explicit "let's analyze carefully" instructions. The two-token cues beat both on MATH-500, AMC 23 and AIME 2024 (Table 4).

Is it just making the model talk longer?

That was my first suspicion. Reasoning cues tend to make answers longer, and longer answers buy test-time compute. With the cue, Olmo's mean response grows from 4.8k to 6.7k tokens.

The paper answers it two ways, and I find both convincing.

Two charts. Left: scatter of MATH-500 pass@1 against mean generated tokens for the 20 candidate openings. The '.\n\n Okay' cue sits near 0.77 at about 6.7k tokens. Openings starting with '$\text' sit at about 13k to 14k tokens but score around 0.3, below no cue. Right: pass@1 against a token budget from 512 to 32k. The cue curve leads the no-cue curve from 1k tokens on and reaches the RL model's accuracy at roughly half the budget the RL model needs.
Length alone does not explain it. Some openings double the response length and still score below no cue; truncated to the same budget, the cue stays ahead from 1k tokens on (paper, Figure 4).

Some openings, such as ones starting a LaTeX \text command, produce responses about twice as long as the cue's and score below no cue at all. And when every response is truncated to the same budget and regraded, the cue beats no cue from 1k tokens onward and reaches the RL model's accuracy at about half its token cap. Longer generation alone is not the story.

There is a cost comparison in Appendix E.2 that I would put in the main text. To match one cued response by majority vote over uncued ones, Olmo-3-7B needs 9 samples and 6.7 times the tokens; across the four models it is 3.6 to 12 times as many tokens. If you serve a base model for math, the opening is the cheapest compute you will ever buy.

The pass@16 numbers point at the same reading from the other side. Olmo-3-7B without a cue already solves 94.0% of MATH-500 problems in at least one of 16 tries; with the cue that becomes 95.4% (Table 6). The cue barely changes which problems the model can solve. It changes how often a single sample lands on the solving behaviour. Keep that in mind for the next section.

What RL was actually doing

This is the part I did not expect to find so cleanly. If two prefilled tokens get you RL's accuracy, did RL learn anything more than to write those two tokens?

The authors compare base and RL models token by token on the same text: give both the same RL-generated history and measure the KL divergence between their next-token distributions at each position. It peaks in the first two positions and is close to zero for the thousands of tokens after. Given the same prefix, the two models want to say nearly the same thing. What changed is the prefix they choose.

Three panels for Olmo-3-7B and Qwen3-14B. Left: token KL divergence between base and RL against token position on a log axis from 0 to 4k. Both curves spike within the first one or two positions, up to about 1.4 for Olmo and 3.3 for Qwen, and are near zero after. Middle: probability trees. Olmo base: '.\n' 0.38, '.\n\n' 0.25 then 'Okay' 0.57. After RL: '.\n\n' 0.66 then 'Okay' 0.98. Qwen base: ' Alright' 0.04 then ',' 1.00. After RL: ' Alright' 0.58. Right: MATH-500 pass@1 over 300 GRPO steps. Cue-forced RL starts near 0.78 for Olmo and 0.85 for Qwen and barely moves; standard RL starts at 0.43 and 0.71 and climbs to meet it.
Where RL changes the policy. Divergence from the base model peaks at the first two output positions, RL turns the cue into the likely opening, and starting RL with the cue already forced gives away little accuracy that standard RL later finds (paper, Figure 3).

RL raises the probability of .\n\nOkay from 0.14 to 0.65 on Olmo-3-7B, and of Alright, from 0.04 to 0.58 on Qwen3-14B. And cue-forced GRPO, where the cue is prefilled and excluded from the loss, starts at the accuracy that standard GRPO reaches after about 100 steps on Olmo and 50 on Qwen, then adds little.

The experiment that convinced me is the swap in Appendix D.3. Take a response, cut it two tokens after the first blank line, and hand the rest to the other model:

Opening fromContinued byMATH-500 pass@1
BaseBase41.7%
BaseRL43.7%
RLRL75.1%
RLBase77.7%

The RL model, forced to continue the base model's opening, is barely better than the base. The base model, continuing the RL model's opening, slightly beats RL. On this setup, almost all of what 300 steps of GRPO bought is in the opening. In 99% of RL rollouts the first blank line is followed by "Okay,".

The toy below is the factorisation from the top of the article with only one knob. Hold the base model's continuations fixed, and move nothing but the probability that it opens with the cue.

move only p(cue | prompt) · base continuation fixedexpected pass@1 42.2%
RL-Zero model 75.0%
base 0.138
after RL 0.65
p(.\n\n Okay | prompt) = 0.14other openings 36.5% · cue 77.9%

A straight line, because the only thing moving is how often the model picks the good branch. At the base model's own 0.138 it lands on the measured 42.2%; it reaches the RL model only as the probability nears 1, which is roughly what the RL model does: 99% of its rollouts follow the first blank line with Okay. The other openings' 36.5% is my arithmetic from the paper's numbers, not a measured value.

With the paper's numbers (42.2% with no cue, 77.9% with the cue forced, 0.138 probability of the cue), every other opening the base model picks averages out to 36.5%; that last figure is my arithmetic, not theirs. At the RL model's 0.65 the line gives about 63%, short of RL's 75.0%. The gap closes as the probability nears 1, which matches the swap table better than the 0.65 does: the 0.65 is the probability of those exact two tokens first, while the RL model also gets to the blank line by other routes and then says "Okay" 99% of the time. The toy is a sketch. The swap experiment is the evidence.

Two more pieces of Appendix D push in the same direction. The rank-64 LoRA update from GRPO, decomposed by SVD at the last MLP output projection, is dominated by two directions that account for 55.9% of its squared Frobenius norm, and both read out in vocabulary space as variants of "Okay" and newline tokens. And the authors skip RL entirely with a closed-form weight edit to that one matrix, ΔW∝GC−1\Delta W \propto G C^{-1}, where GG is the gradient of the cue's log-probability on 10 math prompts and CC is the second moment of the layer's input on ordinary mid-training text. The C−1C^{-1} keeps the change out of directions that ordinary text uses: C4 perplexity moves by 0.04% on Olmo, against 0.41% for RL. Scaled until the cue has probability 0.95, the edited base recovers much of the cued accuracy on MATH-500, AMC 23 and AIME 2024 (Figure 19).

My take: in an R1-Zero setting, on these models, RL with verifiable rewards is mostly a two-token habit plus whatever the base model already does after it. That fits a line of work arguing that RLVR reweights reasoning paths the base model already has instead of growing new ones (Yue et al., 2025), and it fits the observation in the Fireworks piece on distributed RL that more than 98% of weights are bit-identical between adjacent RL checkpoints. It does not say RL is useless. Making a good behaviour the default is worth a lot when nobody will prefill your model for you. It does say that "the model learned to reason during RL" is the wrong sentence for this setting. "The model learned which of its existing habits to start with" is closer. For the other side of that argument, at a scale where RL does seem to open new ground, see Ring-Zero's trillion-parameter run; for how GRPO-style token-level losses relate to the reward they stand in for, see Qwen's first-order view of RL.

Then the chicken

So why "Okay"? The authors went to the data. Olmo 3's mid-training mix contains synthetic reasoning traces, and "Okay" opens 83% of them (82.5% of 148,607 traces, Appendix F.5.1). A correlation, so far. To make it causal they did the expensive thing: edit the training data and train again.

Olmo is the right model for it, because Ai2 released the final pretraining checkpoint, the mid-training mix (allenai/dolma3_dolmino_mix-10B-1025) and the OLMo-core recipe. The authors restart mid-training from the same checkpoint (stage 1, step 1,413,814) on three versions of the 10B-token mix, with the same data order, seed and 4,769 optimiser steps:

The rename asks whether a word with no meaning for maths can become a reasoning cue just by sitting where "Okay" sat. The redirect asks whether you can take "Okay" back and point it at something else.

Olmo-3-7B · 10B mid-training tokens · MATH-500

The released 10B-token mid-training mix, unedited. 83% of its synthetic reasoning traces open with “Okay”.

pass@1 by forced opening (%)
no cue
13.8
.\n\n Okay
31.9
.\n\n Chicken
2.4
.\n\n Hmm
30.4
.\n\n Alright
30.9
force
P(Chicken | prompt, .\n\n) = 2e-7
next:Little0.08McN0.08
Chicken nuggets are sold in sets of 6, 9, or 12. […] Answer: 18
wrong

Paper numbers, not mine: Table 25 for accuracy, Figure 5 for the probabilities, Section F.1.2 for the replies (the first of 32 rollouts on one factoring problem). Hmm and Alright are the controls nobody edited.

Both work. On the base mix, .\n\nChicken scores 2.4% because the model does what any model would do with a paragraph that starts "Chicken": it writes about chicken. The paper's example is a factoring problem, where the base-mix model answers by inventing a chicken-nugget problem and answering that. On the rename mix the same opening scores 37.2%, next to .\n\nOkay's 38.3%, and the model writes "Chicken, so I need to factor the expression…" with a self-check in the middle. Its next-token distribution after "Chicken" changes from "Little" and "McN" to a comma followed by "let" or "so", which is exactly what follows "Okay" in the original.

It still knows what a chicken is. Asked to continue "She roasted the chicken", the rename-mix model writes "and made a salad. She also made a sauce for the chicken." Asked what a chicken is, it says a bird of the family Galliformes. The new association only fires where the old one did: at the start of a paragraph. In the paper's hand-written contexts, it also begins about half of sentence-initial "Chicken"s deliberatively, so the edit is not perfectly contained.

On the redirect mix, .\n\nOkay falls to 0.2%. Prefilled, it now writes a new question ("Okay, what is the first step to solve the equation 2x + 3 = 5x - 4?") and answers that instead. "Chicken", "Hmm" and "Alright", none of them touched by the redirect, all sit at 40%.

Three rows: base mix, rename mix, redirect mix. Each row shows a snippet of edited training text, a bar chart of MATH-500 pass@1 for no cue, Okay and Chicken, and a token probability tree. Base mix: no cue 0.14, Okay 0.32, Chicken 0.02. Rename mix: 0.23, 0.38, 0.37. Redirect mix: 0.27, 0.00, 0.40. In the redirect tree, the tokens after 'Okay,' are 'what' 0.78 and 'which' 0.06 instead of 'let' and 'so'.
The counterfactual runs. Renaming makes Chicken a reasoning cue; redirecting Question: labels to Okay, takes Okay's effect away and points it at writing questions (paper, Figure 5).

Three results in this section surprised me more than the chicken itself.

The first is that renaming did not kill "Okay". After 10B tokens with no "okay" anywhere, the model assigns Okay a probability of about 5×10−75 \times 10^{-7} after the paragraph break, down from 0.15, and yet forcing it still gives 38.3%. The paper's explanation is that the association was already there before mid-training: the cue already helps at the final pretraining checkpoint (Appendix F.6.1), and mid-training widens the gap. You can make a word unlikely without unlearning what follows it. To actually remove the effect, they had to give the word a different job.

The second is "Hmm". In the same 148,607 traces, "Hmm" opens just 13 documents. It scores 30.4% as a cue on the base mix, as good as "Okay" with its 82.5% share. So the association is not a simple count of how often a word opened a trace. My guess is that "Hmm" sits close to "Okay" in whatever the model has learned about deliberative openers from pretraining text, and the trace data sharpens a neighbourhood rather than a single token. The paper does not test that, and I could not check it.

The third is that the edits changed the no-cue baseline too. No-cue accuracy rises from 13.8% on the base mix to 23.3% on the rename mix and 26.7% on the redirect mix, and the edited models start more often with "To" and less with "Answer" (Table 28). The authors say plainly that this accompanies the gain without establishing its cause. It is a reminder that a find-and-replace over 10B tokens is not a surgical intervention.

The edit survives a full-scale run. Applied to the whole 100B-token mix (47,684 steps), the chicken cue reaches 76.9%, .\n\nOkay gets 75.1%, and no-cue rises from 46.7% to 58.2% against the unedited mid-training checkpoint (Table 17). That run is the source of the post's 18% to 77%. It repeats on a second model family too: on SmolLM3-3B, where reasoning traces are only 0.55% of the training tokens, renaming takes the chicken cue from 1.1% to 11.7%, and redirecting drops .\n\nOkay from 19.3% on the base mix to 1.8% (Table 18). The same pattern holds when the edited word is "Alright", renamed to "Duck" (Tables 22 and 23).

The last edit is the one I would show people who think prompts work because of what the words mean. Replace "step by step" and "step-by-step" with "duck duck goose" in the mid-training data, retrain, and put an instruction in the prompt with no prefill at all. "Think duck duck goose" goes from 6.4% to 16.5% on MATH-500; "Think step by step" goes from 12.2% to 14.9%.

Bar chart of MATH-500 pass@1 with no prefill. Base mix: 'Think step by step.' 0.12, 'Think duck duck goose.' 0.06. Rename mix: 'Think step by step.' 0.15, 'Think duck duck goose.' 0.17, the bar decorated with two ducks and a goose.
An instruction is a learned association too: after the data edit, a nonsense phrase works as well as the classic one (paper, Figure 7).

The absolute numbers are small. A 10B-token mid-training run is not a finished model, and both instructions sit near zero on AIME. But the direction is the point: "step by step" helps partly because of where the phrase sat in training documents, and a phrase that sat in the same places would help as much.

Inside the model, and where it stops

The authors also look for the cue in the hidden states. They take layer-24 states from generated responses, find each one's ten nearest neighbours among states computed on training documents, and count which source those neighbours came from, relative to no cue.

Sankey diagram. Left: Olmo-3-7B's 20 candidate openings grouped by MATH-500 accuracy: 63 to 77 percent for Okay variants and bare paragraph breaks, 56.6 percent for '.\n\n To', 26 to 48 percent for a group of 'in', 'and', 'without' openings, 13 to 17 percent for 'Answer' openings, 6.3 percent for '.\n Problem'. Right: bands flow to training sources. The top group flows to reasoning traces, '.\n\n To' to expository math, the Answer group to short Q&A, and Problem to meta-reasoning.
Each cue moves the generated text's hidden states toward a different kind of training document, and the most accurate cues move them toward reasoning traces (paper, Figure 8).

.\n\nOkay moves them toward the synthetic reasoning traces. .\n\nTo moves them toward expository maths and produces textbook-style "First, we need to find…" solutions. Answer moves them toward short Q&A. Across 2,000 responses per opening, checking phrases appear in 97% of "Okay" responses and 5% of "To" responses. The honest exception is .\nProblem, which moves the states toward a meta-reasoning source and scores 6.3%: it restates the task, writes a "fine-grained rationale" as a list, and never computes the answer. Looking like reasoning, in representation space, is not the same as reasoning.

Position matters as well. Insert "Okay" at the start of paragraph 1 or 2 of an answer and accuracy is 69% to 71%; insert it at paragraphs 3 to 6 and accuracy drops to 23% to 30%, below the 38% with no cue at all.

Left: bar chart of MATH-500 pass@1 when 'Okay' is planted at paragraph 1 through 6: about 0.71 and 0.69 for paragraphs 1 and 2, about 0.25, 0.25, 0.24 and 0.30 for paragraphs 3 to 6, against a no-cue line near 0.38. Right: similarity to training reasoning traces over token positions from minus 64 to 128 around the planted Okay. Every insertion spikes at the Okay. For paragraphs 1 and 2 the similarity stays high afterward, near 0.65; for paragraphs 3 to 6 it falls back to the no-cue level within a few tokens.
A cue planted early holds the model in the reasoning-trace region; planted late, it spikes and fades, and accuracy falls below no cue (paper, Figure 9).

Every insertion spikes the similarity to reasoning traces. Only early ones keep it there for the next 128 tokens. The model has already committed to a mode by paragraph three, and a late "Okay" reads as noise in the middle of it.

A comma and a space, and the model complies

Section 4 is a short detour into safety, and it is the part I would flag to anyone running base models behind an API that allows prefill. On XSTest's 450 prompts, Olmo-3-7B-Base with .\n\nOkay, refuses 81.9% of unsafe requests and produces a harmful response 0.2% of the time. Change only the leading characters, to a space instead of a period and blank line, and Okay, refuses 38.3% and produces harmful text 42.4% of the time (Table 35). Same word, same comma.

Left: scatter of refusal of unsafe prompts against answering of safe prompts. The base model with no cue sits near 0.56 refusal. 'I'm sorry' moves it to the top-left corner, refusing almost everything. '.\n\n Okay' moves it up toward the Instruct and Think reference models in the helpful-and-safe quadrant. ' Okay,' and ' How' move it down into the answers-everything quadrant. Right: Sankey bands from each cue to training sources: I'm sorry to refusal text, both forms of Okay to reasoning traces, How to code and science PDFs.
Openings change which requests a base model refuses. The math cue happens to give Olmo the selective refusal of its post-trained variants; a near-identical opening without the paragraph break does the opposite (paper, Figure 10).

The paper reads the selective refusal as the reasoning cue carrying over: the model deliberates about the request before answering. I would not lean on it. The appendix adds that on Qwen3-4B and Qwen3-14B, the paragraph-break "Okay" lowers refusal of unsafe prompts, and the authors call Olmo's benefit specific to Olmo among the models they tried. What generalises is the sensitivity, which is what prefill jailbreaks have exploited for years: whoever controls the first tokens controls a lot of what follows.

The limits, as I read them

The paper is careful about its scope, and I want to be at least as careful.

It is one training regime. Every RL comparison is R1-Zero style: GRPO straight on a base model, no SFT first. The authors' own runs are 300 steps of LoRA at rank 64 with a 4,096-token rollout cap. A multi-stage post-training pipeline with SFT on distilled traces, longer rollouts and full fine-tuning could add things two tokens cannot. The swap experiment, the strongest evidence, is on Olmo-3-7B only.

It depends on the model having the association. Llama-3.1-8B has no effective cue among the openings tested: under the RL-Zero prompt 94% of its rollouts have no extractable answer, and .\n\nOkay reaches 2.0%. Qwen2.5-Math-7B, which already writes solutions by default, gains 2 to 3 points from the best cue and loses 3 from .\n\nOkay. The paper's own account predicts this: the cue is learned, so a model whose data never paired an opener with worked reasoning has nothing to cue. Olmo and SmolLM3 were mid-trained on synthetic reasoning traces with very consistent openers, and the authors note that may make the effect unusually strong.

The data edits are single runs. Each mix is trained once, and the error bars in the tables are standard deviations over rollouts, not over training seeds. With effects of 2.4% to 37.2% that does not worry me for the headline. It does for smaller gaps such as "duck duck goose" at 16.5% against "step by step" at 14.9%.

And the repository does not cover them. sophicle/cues contains the cue search, the evaluation harness and the GRPO training with and without forced cues. The mid-training data edits, the hidden-state neighbour analysis and the weight edit are described in the appendices with their settings but are not in the code. The 10B-token runs need Ai2's checkpoint, mix and recipe, which are public, so they are reproducible in principle. I have not seen anyone do it.

What I would do with it

If I were evaluating a base model on reasoning, I would now report the opening I used, or better, run the label-free search and report both numbers. A base-model score with an uncontrolled first token is partly a measurement of which habit the model happened to start in. The pass@16 numbers say the ability was there; pass@1 says how often you meet it.

If I were planning RL on a base model, I would run the cue search first. It costs GPU-hours, not GPU-weeks, and it tells you how much of the RL gain you are about to buy is already sitting in the base model's second-likeliest opening. Cue-forced RL started where standard RL got to after 100 steps on Olmo.

And if I were building a mid-training mix with synthetic traces, I would treat the trace format as a design decision. Whatever word the generator likes to start with becomes a switch in the model. Here it was "Okay", and it happened to be a useful one. The paper's own closing point is that such associations can be put in on purpose, and can also slip in by accident.

How I checked

I read the arXiv HTML version of the paper in full, including the appendices, and took every number in this article from its tables and text: Table 6 and Table 7 for base and RL results with pass@16, Table 3 for random openings, Table 4 for GEPA, Table 14 for the swap, Tables 16 to 28 for the data edits, Table 35 for the safety numbers, and Appendix D for the RL and weight-edit details. Where the figure captions and the running text disagree on a figure number (the HTML's cross-references are off by a few in Sections 2 and 3), I used the caption's number. The figures here are the paper's own, downloaded from arXiv; the SVG ones were rendered to PNG in headless Chromium.

I shallow-cloned sophicle/cues and read the cue search (cues/beam.py, cues/search.py, cues/agree.py), the prompts and cue prefill (cues/prompts.py), the model registry with the pinned revisions and selected cues (cues/models.py), and the GRPO trainer (cues/rl/train.py). That is where the beam-search defaults that differ from the paper's description come from. I did not run any of it: there are no GPUs on this machine, and the brief for this site is not to execute third-party code. Nothing here is my measurement. The one number that is my own arithmetic, the 36.5% average accuracy of the base model's other openings, is labelled where it appears. The project's blog page sat behind a browser challenge I could not pass, so I relied on the paper, the thread and the code.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Token cues: how 'chicken' made a base model reason as well as RL", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026chickenreasoningcues,
  author = {Satyajit Ghana},
  title  = {Token cues: how 'chicken' made a base model reason as well as RL},
  url    = {https://ai.thesatyajit.com/articles/chicken-reasoning-cues},
  year   = {2026}
}
share