# A $80 paper rewriter: what the fine-tune bought, and what the judge decided

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/paper-rewriter-finetune
> date: 2026-10-06
> tags: fine-tuning, distillation, evaluation, small-models, lora, synthetic-data, qwen, llm

On 5 October, Maxime Rivest [posted](https://x.com/MaximeRivest/status/2106916467077750913):

> I never thought that with only 5000 training examples, 4 hours of training and 80\$ I could make a 4b model beat gpt 5.6 terra.

One reply asked the right question: "pairwise wins can just mean the 4b learned what the grader likes." Unusually, it can be answered. Within a day he had published both models ([4B](https://huggingface.co/maximerivest/qwen3.5-4b-paper-rewriter), [9B](https://huggingface.co/maximerivest/qwen3.5-9b-paper-rewriter)), the [training data](https://huggingface.co/datasets/maximerivest/paper-rewrites-for-young-readers), a [write-up](https://x.com/MaximeRivest/status/2107149260080685248), a reading app at [sciencemadereadable.com](https://sciencemadereadable.com), and the [code](https://github.com/MaximeRivest/sciencemadereadable), including the raw judge verdicts.

So I read the cards against the weights, re-tokenized the dataset, and recomputed the leaderboard from his result files. The short version:

| | what I found |
|---|---|
| Task | rewrite a whole paper, section by section, so a 12–14-year-old can read it; not a summary |
| Data | 4,416 PubMed Central papers → 20,394 training conversations, **189,983,131 tokens** (measured) |
| Teachers | Claude Opus 5.5 (13,877 answers) and GPT-6 Astra (6,997) (reported, dataset card and repo) |
| 4B | LoRA rank 32, one pass, one RTX 3090 at home, 38 h (reported); the weight delta is rank 32 (measured) |
| 9B | all weights, one pass, one B300, 6 h, "roughly \$80" (reported); the delta is full rank (measured) |
| "beat Terra" | aggregate pairwise win rate, 40% vs 35%, judged by Opus (reported, reproduced) |
| 4B vs Terra directly | **6 wins, 7 losses**, 2 no-clear-winner (reported, reproduced) |
| Same pairings, Astra judging | Terra beats the 9B 15-0 and the 4B 14-0 (measured from his `results-astra.json`) |

The recipe is real and worth copying. The headline is a judge's opinion, and the second judge disagrees.

## The task: rewrite, don't summarise

The card states the job in one line: rewrite a research paper so a curious 12–14-year-old can read it, "the same paper, in the authors' own voice, keeping every result, number and uncertainty, in plain words. It rewrites, it does not summarise."

That framing matters for everything downstream. A summary can drop a hedge and still be a good summary. A rewrite that drops a hedge is wrong. So the task has two halves that pull against each other: readability, which a small model picks up quickly, and fidelity, which it does not.

The pipeline makes two kinds of call per paper. The **opening** call sees the whole paper and rewrites the title, abstract, first introduction paragraph and conclusion. Then each of up to four **section** calls (`introduction_rest`, `methods`, `results`, `discussion`) rewrites one section, given the rewritten opening so the voice and the plain terms carry over. Every call also gets a `reference_glossary`: definitions of the passage's technical terms, pulled from an offline copy of Wikipedia and Wiktionary. The shipped `rewrite.py` reproduces those prompts character for character; the system prompt literally begins `Function: rewrite_opening_student`.

Here is what the 4B produced on a grassland-management paper he posted before the claim:

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig3.jpg"
  alt="Two-column table. Left, the original introduction of a paper about plateau pika control and grassland fencing on the Qinghai-Tibet Plateau, in dense academic prose. Right, the 4B model's rewrite in plain sentences, keeping figures such as 4 million km2, more than 25 years, 78% of degraded grasslands, \$300,000 to \$800,000 and 38,000 tons of rodenticide."
  caption="Original (left) and the fine-tuned 4B's rewrite (right) of an ecology paper's introduction. The numbers survive; the jargon is unpacked. Note the phrase 'Brandt's vole, another small rodent': pikas are not rodents, and the repo's own design notes name 'the pika is a rodent' error. (Maxime Rivest, X post of 4 October 2026, image 1 of 3.)"
/>

It reads well, and every number I checked against the left column survives. It also makes the error that is hardest to catch: "Brandt's vole, another small rodent" quietly calls a pika a rodent. Pikas are lagomorphs, cousins of rabbits. The original never says otherwise because it never needed to. This is the failure the card describes under limitations ("two similar technical terms confused") and that the repo's `training/dataset_v3.md` names outright: the model "fills the gap from memory (the 'pika is a rodent' error)". I found it in the showcase image, not by hunting.

## From 5,000 papers to 190 million tokens

The "5000 training examples" in the post are papers, not examples. The write-up says he started with 5,000 open-access ecology and environmental-science papers; frontier models rewrote 4,416 of them.

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig1.jpg"
  alt="Pipeline diagram: 5,000 open papers from PubMed Central under CC BY, then frontier models rewrite them (4,416 papers to 21,069 examples), then small Qwen 3.5 models learn (0.8B, 4B and 9B, one pass), then Opus judges every rewrite (3 held-out papers, 15 parts, 4 scores out of 10). A reference glossary from offline Wikipedia feeds both the writers and the small models."
  caption="The whole recipe: frontier teachers write the rewrites, small Qwen 3.5 students copy them, an Opus judge scores everyone. (Maxime Rivest, 'Even you can now fine-tune an LLM', X article, figure 1.)"
/>

I downloaded the dataset and counted. The train split has 20,874 conversations from 4,376 papers; validation has 195 from 40 more. By call type: 4,376 openings, 4,376 methods, 4,376 results, 4,353 introductions and 3,393 discussions (measured). The writer labels are anonymised as `writer_a` and `writer_b`, but the repo's `training/hf/make_dataset.py` maps them: `{"opus": "writer_a", "astra": "writer_b"}`. So two-thirds of what the students learned to imitate, 13,877 of 20,874 training answers, is Claude Opus 5.5 (measured). Remember that when we get to the judge.

With the model's own tokenizer, the train split is 204.5M tokens. Drop the conversations longer than 24,576 tokens, as the card says training did, and **exactly 480** go, leaving 20,394 conversations and **189,983,131 tokens** (measured). That is the card's "190M tokens". The average conversation is 9,316 tokens; the median is 8,212. Only 51.1M of those tokens, 27%, are the assistant's answer, the part the loss is computed on (measured). That number is worth having, because his own training-loss chart puts "answer tokens learned" on its x-axis, and the 9B curve ends just past 50M. It matches.

The step that mattered most happened before any of this was generated. He iterated the teacher prompts eight times against his rubric, and the judge's score rose by 1.5 to 2.8 points for every frontier model:

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig4.jpg"
  alt="Dumbbell chart: Opus rises from 5.6 to 8.5 (+2.8), Astra from 5.8 to 8.2 (+2.3), Luna from 6.0 to 7.5 (+1.5) between first prompts (v1) and refined prompts (v8). Footnote: part of the gap is a change of goal, since v1 also corrected the paper and allowed lists."
  caption="Prompt refinement moved the teachers more than any later training choice moved the students. The footnote is honest: part of the gap is a change of goal. (Maxime Rivest, X article, figure 2.)"
/>

His own sentence is the best one in the write-up: "whatever the frontier model does, your small model will copy, mistakes included." The student's ceiling is the teacher's output, so the cheapest point of leverage is the teacher's prompt. Hold on to a second implication too: those eight rounds were scored by the same Opus judge that later grades the students.

## Training: what the weights say

The two cards describe two different runs.

| | 4B | 9B |
|---|---|---|
| Method | LoRA rank 32, alpha 64, every linear layer, lr 2e-4 | all weights, lr 1e-5 |
| Hardware | one RTX 3090, at home | one B300, rented |
| Time | 38 h | 6 h |
| Passes | one | one |
| Released as | merged into full bf16 weights | full bf16 weights |

(Reported, model cards and `training/train.py`.) A merged LoRA looks like any other checkpoint, so I checked. I pulled single tensors out of the fine-tuned `model.safetensors` and the base `Qwen/Qwen3.5-4B` shards with HTTP range requests, subtracted, and took the singular values of the difference.

For the 4B, layer 10's `mlp.down_proj` (2,560 × 9,216) has a delta whose 32nd singular value is 0.2564 and whose 33rd is 0.0042; the top 32 directions hold 99.95% of the change, and what is left is bf16 rounding (measured). Layer 11's `q_proj` shows the same cliff at exactly 32. That is a rank-32 LoRA, merged. The relative change in the weight is 7.1%. For the 9B, the same `down_proj` (4,096 × 12,288) has no cliff at all: the top 32 directions hold 10.9% of the change, and the relative change is 2.9% (measured). Full fine-tuning, as stated. The cards describe the weights.

The hardware claim is checkable from arithmetic alone. The write-up text says the 9B ran on "one rented H100, roughly \$80", but its own loss chart labels the run "full + glossary (Nebius B300)", and so does the card. The training script keeps fp32 master weights and AdamW state for a full fine-tune: 4 bytes of weight, 4 of gradient and 8 of optimizer state per parameter. At 8.954B parameters (measured, from the safetensors header) that is 143 GB before a single activation (reasoned). It does not fit on an 80 GB H100. It fits on a B300.

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig5.jpg"
  alt="Two panels. Top: training loss against answer tokens learned (0 to 60 million) for Qwen3.5 0.8B, 2B, 4B and 9B runs, the 9B run labelled 'full + glossary (Nebius B300)'. Bottom: best validation loss against parameter count, falling from about 0.57 at 0.8B to about 0.35 at 9B."
  caption="Training loss over answer tokens, and best validation loss against model size. The 9B legend entry names the B300. (Maxime Rivest, X article, figure 4.)"
/>

## Where the \$80 and the 4 hours come from

Now line the tweet up against the cards. "5000 training examples" is 4,416 papers, 20,394 conversations. "4 hours" is closest to the 9B's 6. "\$80" is the 9B's bill in the write-up. The 4B never cost \$80: it trained for 38 hours on a card he already owned. The tweet compresses two runs into one sentence and attaches the cheap run's model size to the rented run's price.

Is \$80 for 6 hours of one B300 plausible? That is \$13.33 an hour (reasoned). A third-party price list puts Nebius's on-demand B300 at \$7.85 an hour before 1 October and \$9.50 after (reported by Spheron, not Nebius). At those rates 6 hours is \$47 to \$57 (reasoned), so \$80 covers the training plus the setup, evaluation and idle time an agent-managed rental accrues. That is a believable bill, and an unverified one.

The more useful comparison is the line item the tweet leaves out. The write-up says having frontier models write the 4,416 papers' rewrites "would have cost about \$4,600" at pay-per-use prices (reported). He says elsewhere the big models ran through his Claude and ChatGPT logins, so his actual cash cost may have been lower. Either way, at list prices the GPU is about 1% of the project (reasoned). Throughput is easy to calibrate from his two runs: the 9B pushed 8,796 tokens a second through the B300, the 4B LoRA 1,389 through the 3090 (reasoned, tokens over wall-clock). The calculator below runs on those two rates, scaled linearly by parameter count, which is the 6ND rule of thumb and is optimistic for small models that leave a big GPU idle. The default home-3090 price in it is my assumption, not his.

<FinetuneCost />

At his defaults that is 6.0 hours and \$47 of B300 against about \$4,600 of teacher rewrites, about \$1.04 a paper (reasoned). The GPU line is the rounding error.

And you may not need most of it. His checkpoint scores say the judge stopped improving long before the loss did:

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig6.jpg"
  alt="Validation loss against share of the training pass for two runs. The 4B first run (no glossary) falls from 0.73 to 0.40, and scores judge 6.9 both at 16% of the pass and at the end. The 9B falls from 0.66 to 0.35 and scores judge 7.9 at 33% of the pass and 7.8 at the end."
  caption="Validation loss keeps falling to the end of the pass; the judge score does not. The 9B checkpoint a third of the way in scored 7.9, the final one 7.8. (Maxime Rivest, X article, figure 5.)"
/>

In his words: "training on more then about 1000 examples bought nothing I could measure." Two caveats. "Measure" here is a 15-part judge score with an error bar of a few tenths, so it only sees large effects. And the still-falling loss means the student keeps moving toward the teacher token by token, possibly on the fidelity the rubric scores too coarsely to notice. The direction is still clear: score checkpoints with your rubric, not your loss curve.

## What the scores say before anyone plays head to head

The card's absolute scores, Opus judging each rewrite 0–10 on four aspects over the same 15 parts, are the cleaner view:

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig7.jpg"
  alt="Four bar panels: Faithful and exact, Understandable, Pleasant to read, Structure and voice. Before and after training for 0.8B, 4B and 9B against dashed lines for Opus, Astra and Luna. On faithfulness the 4B goes 2.7 to 6.1 and the 9B 4.3 to 7.1, below Luna 7.2, Opus 8.3 and Astra 8.5. On structure the 4B reaches 8.5 and the 9B 8.6, level with the frontier models."
  caption="Training closes the gap on clarity, readability and structure; faithfulness is the last mile. (Maxime Rivest, X article, figure 6.)"
/>

The 4B averages 7.2 and the 9B 7.8; Luna and Terra both score 7.5 (reported). On faithfulness the 4B is at 6.1 against Luna's 7.2, and the card counts about 11.3 serious problems per paper for the 4B, 5.7 for the 9B, 9.0 for Luna and 0.7 for Opus (reported). So on the absolute rubric the 4B is **below** both GPT models, and makes more serious errors per paper than either GPT model the card lists. The 9B is a little above them. Structure and voice, the part a student copies first, is where both students match the teachers.

The 4B card breaks the same scores out part by part. Against GPT-6 Luna, on the average of the four aspects, the 4B is better on 3 parts, level on 3 and worse on 9 (reported):

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig9.jpg"
  alt="Two panels of stacked bars, better, same or worse on 15 parts, for the 4B against four models. Average of the four aspects: vs Claude Opus 5.5 1 better, 1 same, 13 worse; vs GPT-6 Astra 2 same, 13 worse; vs GPT-6 Luna 3 better, 3 same, 9 worse; vs untrained Qwen3.5-4B 15 better. Faithful and exact: vs Luna 2 better, 3 same, 10 worse."
  caption="The 4B part by part on the absolute rubric. It loses to Luna 9 parts to 3, the opposite of the pairwise chart's ordering. (qwen3.5-4b-paper-rewriter model card, winrate.png.)"
/>

## How "beat Terra" was measured

The tweet's chart is not those scores. It is a round robin. Nine writers (six frontier models: Opus, Sonnet, Astra, Sol, Terra 5.6 and Luna; and the three students), every pair, on 15 parts of 3 held-out papers: 36 pairs × 15 parts = 540 pairings. Opus sees the original part and two rewrites labelled only A and B, and says which "a curious 14-year-old would find more pleasant and easier to follow, given that it must stay faithful", or tie. Each pairing is asked in both orders, and a win counts only if it holds both ways; otherwise it is "no clear winner" (reported, `training/head_to_head.py`). That is a careful design. Of the 540 pairings, 64 flipped with presentation order and became no-clear-winners (measured from `results.json`); a single-order judge would have scored those as wins for whichever went first.

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig2.jpg"
  alt="Horizontal stacked bars of wins, no clear winner and losses for nine writers, each facing 8 others on 15 parts (120 matchups). Win rates: Claude Opus 5.5 88%, GPT-6 Astra 75%, Claude Sonnet 5.5 67%, GPT-6 Sol 59%, Our 9B 55%, Our 4B 40%, GPT-5.6 Terra 35%, GPT-6 Luna 29%, Our 0.8B 2%."
  caption="The chart behind the claim. 'Win rate' is wins plus half the no-clear-winners, over 120 matchups per writer, against all eight others. (Maxime Rivest, X post of 5 October 2026.)"
/>

The win rate column is (wins + ½ no-clear-winner) / 120. For the 4B: 40 wins, 64 losses, 16 no-clear-winners, so 48/120 = 40.0%. For Terra: 34, 69, 17, so 35.4% (measured, recomputed from his `results.json`; every bar matches). The 4B sits above Terra on this chart. But this number averages over **all eight opponents**. The 4B does a little better than Terra against Opus (1–12 vs 0–14), Sonnet (3–11 vs 1–11), Sol (5–8 vs 1–11) and the 9B (4–8 vs 4–11), and that is where its edge comes from. When the two actually meet:

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig8.jpg"
  alt="Nine by nine heat map of row-model win rates against column models, 15 parts each. Our 9B against Terra 5.6: 73%, 11-4. Our 4B against Terra 5.6: 47%, 6-7. Terra against Our 4B: 53%, 7-6. Opus beats every model; Our 0.8B loses to every model."
  caption="Who beats whom. Read across a row. The 4B's row against Terra 5.6 reads 47%, 6–7. (Maxime Rivest, X article, figure 7.)"
/>

**The 4B loses to Terra head to head, 6 to 7, with 2 parts undecided** (reported, and reproduced from the result file). The 9B beats Terra 11 to 4, and that is the claim the write-up makes, correctly. The tweet's version, a 4B beating Terra, is true only of a leaderboard average, and false of the matchup the sentence describes.

How much is 11 to 4 worth? If the two were equally good, eleven or more wins out of fifteen decisive parts happens with probability 1,941/32,768, about 6% (reasoned, a one-sided sign test). And the 15 parts are not 15 independent samples: they come from 3 papers, and a judge that prefers one writer's handling of a paper's jargon will prefer it in all five of that paper's parts. His card says the same thing in plainer words: "three papers is a small test, so gaps under about half a point may be luck."

## Swap the judge

Here is the experiment the replies asked for, and he already ran it. `head_to_head.py` has a `--judge astra` flag, and the repo commits both result files: the same 540 pairings, the same rewrites, the same both-orders rule, with GPT-6 Astra judging instead of Claude Opus. He mentioned it in a reply: "astra would score these differently, but worst, in my opinion." Differently is an understatement.

<JudgeFlip />

Under Opus, the ranking is Opus 88%, Astra 75%, Sonnet 67%, Sol 59%, 9B 55%, 4B 40%, Terra 35%, Luna 29%, 0.8B 2%. Under Astra it is Astra 93%, **Terra 78%**, Luna 63%, Sonnet and Sol 61%, Opus 43%, 9B 30%, 4B 20%, 0.8B 1% (measured, both recomputed from the committed files). Each judge puts its own lab's flagship first. Terra goes from second-to-last among the frontier models to second overall. The students go from mid-table to the bottom of everything that can write. Head to head under Astra, Terra beats the 9B 15–0 and the 4B 14–0, with one undecided.

This is the textbook shape of self-preference in LLM-as-judge evaluation, and here it has a mechanism you can point at. The students were trained on answers that are two-thirds Opus. The teacher prompts were refined for eight rounds against the Opus judge. Then the Opus judge scored the students. Every stage of the loop pulls toward one model's taste, and the evaluation inherits it. That does not make the Opus verdicts wrong; Rivest says he read the verdicts against the papers until he trusted them, and he may simply be right that Opus is the better reader. It does mean "beats Terra" is a statement about Opus's preferences, and the evidence that those preferences track a 14-year-old's is one person's spot checks.

The absolute-score view is less exposed, because it scores faithfulness separately from taste, and there the 9B is still a reasonable result: 7.8 on average against 7.5 for the GPT models, 7.1 on faithfulness against Luna's 7.2, at a fraction of the cost. It is the pairwise "who reads more pleasantly" question, a preference by construction, that flips with the judge.

## The economics that do generalise

The cost argument at inference time does not depend on any judge. A rented GPU costs the same per hour whether it works or not, so the student wins only if it is kept busy:

<Figure
  src="https://ai.thesatyajit.com/articles/paper-rewriter-finetune/fig10.jpg"
  alt="Log-log line chart of price per paper against average papers per hour for the 9B on a rented RTX PRO 6000. The line crosses Opus's \$0.97 per paper at 1.8 papers per hour and Luna's \$0.011 per paper at 168 papers per hour; one GPU tops out at 392 papers per hour, after which the line saw-tooths as GPUs are added."
  caption="Break-even speed is the GPU's price per hour divided by the frontier model's price per paper. (Maxime Rivest, X article, figure 9.)"
/>

His formula is the right one: break-even speed = GPU price per hour ÷ frontier price per paper. At \$1.80 an hour for an RTX PRO 6000 and \$0.01072 per paper for Luna, that is 168 papers an hour (reported, and the arithmetic checks). The 9B tops out at 392 papers an hour on that card (reported), so it has to run at 43% of capacity, around the clock, just to match Luna's price (reasoned). Against Opus at \$0.97 a paper it wins at 1.8 papers an hour. Against Terra, which the write-up prices at about \$250 per 1,000 papers, it wins almost immediately.

The up-front money pays back on the same logic. About \$4,680 of teacher rewrites and training, saving \$0.00614 per paper against Luna (\$10.72 vs \$4.58 per 1,000), takes about 762,000 papers to recover (reasoned). Against Opus it is about 4,800 papers. At his stated scale, about 100 million papers, the 9B costs about \$458,000 against \$1.07 million for Luna and \$97.5 million for Opus, and needs about 255,000 GPU-hours on that card (reasoned, at his per-paper prices). The small model is not a toy at that scale; it is the only line on the chart that fits a grant.

## What generalises, and where it breaks

What transfers from this, cleanly:

- **Narrow-task distillation is cheap to train and expensive to feed.** The GPU bill is tens of dollars; the teacher bill is thousands at list prices. Budget the data, not the run. Read the provider terms too: he flags that "not every provider lets you train on its outputs."
- **The teacher's prompt is the highest-leverage step.** Eight rounds moved the teachers by 1.5 to 2.8 points. No training choice moved the students that much.
- **Score checkpoints with the task's rubric.** The loss kept falling; the judge flattened at a third of a pass, about a thousand examples' worth of signal in his telling.
- **Students copy form first and facts last.** Structure reached 8.5–8.6 in both students. Faithfulness is where the 4B sits a full point below Luna, and where the pika became a rodent. The fix in his `dataset_v3.md` plan is the right shape: make every fact traceable to the paper, the glossary or school-level knowledge, and train the model to emit a marker instead of guessing.

What does not transfer, or needs another judge before it does:

- **A pairwise win rate is an average over a field.** Report the head-to-head the sentence claims. Here the leaderboard says 4B over Terra and the matchup says 6–7.
- **One judge from the same family as the main teacher is not an evaluation.** It is a measure of how well the student imitates that family. The cheapest correction already exists in his repo: run the second judge and publish both, which he did. The next one is people. Fifteen rewritten parts read by a handful of actual 14-year-olds, and checked against the paper by someone who knows that a pika is not a rodent, would settle more than another 540 model verdicts.
- **Three papers is a demo.** He says so on the card.

None of this makes the project less useful. Two open models that rewrite a paper in 9 to 15 seconds on one H100, an open dataset of 21,069 examples, and a two-judge evaluation with the raw verdicts committed is more than most releases ship. The tweet was the least careful part of a careful piece of work.

## Related

The same "fine-tune a small model on your own narrow task" argument, from the tool-calling side, is [Cactus Needle 2: the interesting number is what happens after you fine-tune it](/articles/needle-finetune), and its follow-up [Cactus Needle 3: the free fine-tune deletes the confidence head](/articles/needle-3-finetune). For a distillation whose headline came apart under its own card, see [Qwen3.8-2B-Distill: the filter wrote most of the headline](/articles/qwen3-8-2b-distill). [BTL-3: a rank-32 LoRA that turns Qwen3.6-27B into a tool-use agent](/articles/btl-3) is the same LoRA rank applied to agents. And if the prompt-refinement step is the part you want to automate, [GEPA: optimize anything you can score and describe](/articles/gepa-optimize-anything) is a method for exactly that, with the same caveat: it optimizes toward whatever scores it.
