How articles are scored
Every article here carries an editorial rating: eight questions a reader would ask, each answered 0 to 3 by hand. This is what each answer means, how answers become a tier, and how the ratings are kept honest.
On the articles list each card shows a tier and a couple of highlights in plain words. On an article, the Why read this panel adds the one-line reason it is worth your time, and its How this was scored disclosure opens the eight answers behind it.
A rating judges the subject and this page as you would read it. A landmark paper explained badly should not score well, and a modest tool taken apart carefully can.
The eight questions
Each is answered 0 to 3. The anchors are meant to be stingy: a 2 is a good page, and a 3 is rare.
1Is it new?
novelty · counts once- 0 of 3
- Repackaging or news of a known thing
- 1 of 3
- An incremental tweak
- 2 of 3
- A real new idea, method or capability
- 3 of 3
- Changes how the field does something
Shown on a card as “Genuinely new idea” at 3 and “A new technique” at 2.
2Can I trust it?
verification · counts a little more- 0 of 3
- Restates claims
- 1 of 3
- Spot-checks a few numbers
- 2 of 3
- Measures key facts from files, code or configs
- 3 of 3
- Reproduces the headline result, or shows from primary files it is wrong
Shown on a card as “Checked against the source” at 3 and “Facts measured from source” at 2. A 3 on both trust and “only here?” is said once, as “Original, source-checked analysis”.
3Can I run it?
runnable · counts a little less- 0 of 3
- Closed, nothing to run
- 1 of 3
- API-only, gated or restrictive licence
- 2 of 3
- Open code or weights with real limits
- 3 of 3
- Open, permissive, runs on reader hardware with instructions
Shown on a card as “Open and runnable” at 3 and “Open code or weights” at 2, or by the hardware it runs on when that is a browser, phone, laptop or consumer GPU.
4Will I understand it?
explains · counts once- 0 of 3
- Describes, no mechanism
- 1 of 3
- Partial mechanism
- 2 of 3
- Mechanism from first principles with figures
- 3 of 3
- Mechanism carried by interactives built from real code or data
Shown on a card as “Interactive explanations” at 3 and “Explained from first principles” at 2.
5Can I act on it?
takeaway · counts once- 0 of 3
- Nothing to act on
- 1 of 3
- General advice
- 2 of 3
- A concrete recipe, numbers or comparison
- 3 of 3
- A decision guide a practitioner can follow today
Shown on a card as “A guide you can follow today” at 3 and “Concrete numbers to act on” at 2.
6Will it last?
durability · counts a little less- 0 of 3
- News that expires in weeks
- 1 of 3
- Relevant for months
- 2 of 3
- A reference for a year or more
- 3 of 3
- Evergreen fundamentals
Shown on a card as “Evergreen reference” at 3 and “A lasting reference” at 2.
7Does it affect many?
reach · counts a little less- 0 of 3
- A niche paper or one-off repo
- 1 of 3
- A specialist community
- 2 of 3
- A widely used model, tool or lab release
- 3 of 3
- Something most practitioners touch
Shown on a card as “Something most practitioners use” at 3 and “Widely used” at 2.
8Only here?
unique · counts once- 0 of 3
- Same as the press release
- 1 of 3
- Some original analysis
- 2 of 3
- A teardown or measurement few others did
- 3 of 3
- The only place this analysis exists
Shown on a card as “Analysis found nowhere else” at 3 and “Original analysis” at 2. A 3 on both trust and “only here?” is said once, as “Original, source-checked analysis”.
From answers to a score
The eight answers are combined into one number from 0 to 100. Trust counts a little more than the rest; runnability, durability and reach a little less. On top of that, a page earns up to 6 points for evidence it actually carries rather than claims: the paper's own figures, interactive explanations, numbers measured for the article, model and repository cards, and an explainer film. Those are counted from the page, never entered by hand, and each has diminishing returns, so padding a page does not pay.
The number itself stays in the detail view and the JSON. What a card shows is the tier, because a rank among peers says more than a raw score.
Tiers are percentiles
Every rated article is ranked by score, and the tiers take fixed shares of that ranking. However the scores drift, only a tenth of the articles can be Essential. Ties break on the most recently updated, then on the address, so the same commit always gives the same tiers.
| Tier | Band | Articles | Scores now |
|---|---|---|---|
| Essential | Top 10% | 44 | 76–91 |
| High | Top 30% | 90 | 68–76 |
| Notable | Top 60% | 133 | 58–68 |
| Solid | Top 85% | 111 | 44–58 |
| Niche | Last 15% | 67 | 11–44 |
445 rated articles. Drafts take no place in the ranking.
Ways to sort
The list can be ordered for what you want from it. Each order reads only the answers that matter to it.
- Must read
- The best pages here, by overall score.
- Run it yourself
- Open, runnable, with steps and numbers to act on.
- Learn the idea
- Clear mechanisms that will still be true next year.
- What's new
- New ideas with wide reach, freshest first.
- Deep dives
- Measured, verified, analysis you won't find elsewhere.
- Newest
- Everything, most recently updated first.
Keeping it honest
These are editorial judgements, not measurements, made by the same editor and agent crew that writes the pages. Four things keep them in check:
- Stingy anchors. The expected average on every question is about 1.5. Once 50 articles are rated, the site's checks (
validate:ratings) fail the build if any question's average across the corpus passes 2.2. An earlier, looser rating ended with most articles marked High; this one is built so that cannot happen quietly. - Tiers by rank. Shares are fixed, so a generous week cannot promote everything.
- Evidence is counted, not claimed. The bonus comes from what the page carries, read off the page at build time.
- Everything is public. Every answer, score, rank and fact is in the articles JSON and in llms.txt, and every rating sits in its article's source next to the one-line reason for it.