# Impeccable: a design vocabulary for agents is mostly a linter with opinions

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/impeccable-design-skill
> date: 2026-10-07
> tags: agents, agentic-coding, developer-tools, generative-ui

Impeccable's home page has one line of pitch: "The missing design vocabulary for agents. Turn AI slop into interfaces you're proud to ship." Under it is a card. You can drag a slider across it and watch a beige release-planning card with a purple left stripe become a plain white one. The site's star badge reads 76k. I wanted to know what a "design vocabulary" for an agent actually is when you open the box. Is it a list of things not to do, a set of principles, or a tool that measures something?

I expected a long prompt, like most skills I have read here: [figures4papers](/articles/figures4papers) is a house style written as a SKILL.md, and the [one-shot launch video skills](/articles/one-shot-launch-videos) are instructions handed to a framework. What I found is a long prompt and a linter. The linter is larger than everything else in the repository put together, and it holds most of what is new.

<RepoCard repo="pbakaus/impeccable" note="Read at d98b0be (7 October 2026, the depth-1 clone's only commit). Apache-2.0. npm package impeccable 4.1.0; plugin manifest 4.5.0; engine binary 0.1.11." />

## What arrives in your agent

The product is one skill. `skill/SKILL.src.md` is 90 lines. It opens by giving the model a job title:

> "This skill gives you the tools and permission to create design that earns to be called out-of-distribution craft: Whereas before, your design work would have been safe, timid and measured, you now approach every design task as an award-winning design director" (`skill/SKILL.src.md:12`)

The rest of the file is a router. A table maps 24 commands to reference files: `shape`, `critique`, `audit`, `polish`, `bolder`, `quieter`, `distill`, `harden`, `clarify`, `typeset`, `layout`, `colorize`, `animate`, `live` and the rest. One of the 24, `craft`, is a deprecated alias, and the repo's own `CLAUDE.md` counts 23. The table points into `skill/reference/`, which holds 41 Markdown files and about 60,000 words. An agent loads only the file for the command it runs, so the weight is paid per request. It still adds up: `new-work.md` alone is 9,389 words, which is more than this article.

Three instructions in `SKILL.md` hold the rest together. First, run the engine's `context` verb once per session; it loads `PRODUCT.md`, `DESIGN.md` and a per-surface brief (`:21`). Second, load the command's reference. Third, read `craft-floor.md` "immediately before any UI edit, including small refinements" (`:23`). That last file is where most of the vocabulary actually lives, and I come back to it below.

Packaging is a build step, not a hand-maintained fork per tool. `scripts/build.js` transforms the one source tree into a copy per harness, driven by `scripts/lib/transformers/providers.js`: `.claude/skills/`, `.cursor/skills/`, `.agents/skills/` for Codex, `.github/skills/` for Copilot, `.gemini/`, `.opencode/`, `.grok/`, `.hermes/`, `.dsh/` for DeepSeek Harness, Kiro, Pi, Qoder, Trae (two locales), Rovo Dev, Mistral Vibe, Veto and Antigravity. The transform substitutes the script path and command prefix, keeps or drops provider blocks such as `<codex>…</codex>`, and strips 151 distinct `<!-- rule:… -->` markers from the text (`scripts/lib/utils.js:589`). I found nothing in the public repo that reads those markers, so I assume the private evaluation harness does. All the generated copies are committed. They account for 1,178 of the repository's 3,897 files and 43.7 MB of its 85 MB working tree.

Each copy ships a 206-line POSIX launcher, `scripts/impeccable`. On first run it downloads the engine for your platform from the GitHub release into `~/.impeccable/bin/`, and it refuses to run a binary whose `.sha256` sidecar does not match. That check catches a corrupted download. It does not prove who built the binary, since the sidecar comes from the same host. Where a harness has hooks, the installer wires them up. Claude Code gets `PostToolUse` on `Edit|Write` and a `Stop` deep pass (`plugin/hooks/hooks.json`). Cursor gets a `preToolUse` gate that denies a bad write before it lands. Codex needs `/hooks` approval after every update that changes the hook. [DeepSeek Harness](/articles/deepseek-harness) gets the skill but no hook, because its hooks are in-process plugins rather than files on disk.

## Three kinds of vocabulary

Read in full, the skill text sorts into three piles of very different size.

The first pile is principles, and it is small. `SKILL.md` names four visitor modes: Persuade, Operate, Read and Experience. It says the mode follows the surface, not the product, so "a tool's landing page is still Persuade" (`:43`). It also says "the brief wins": "Honor pinned aesthetics, eras, materials, fonts, and palettes even when they conflict with a saturated-pattern warning. Redirecting a clear brief toward your taste is failure" (`:29`). The mode files add sensible defaults. A Read surface keeps "a reading face, real contrast, a real measure, nothing performing behind the text" (`mode-read.md:7`), and Operate prefers one family and a 1.125 to 1.2 type scale (`operate.md:13-15`). None of this is new to a designer. It is new to a model that has never been told which job the page is doing.

The second pile is refusals, and it is the heart of the product. `craft-floor.md` is 52 lines with a "Verify" list and a "Refuse" list. Most of the Refuse list says "these are the category's defaults, not bans" (`:21`). One item is phrased harder than any other line in the skill:

> "A kicker or eyebrow above a heading. This one is a ban, not a default: no brief earns it back." (`skill/reference/craft-floor.md:27`)

The rest of the list is the home page's before/after cards in prose. It covers a coloured `border-left` above 1px (`:35`), gradient text (`:33`), nested cards, "always wrong" (`:25`), monospace "as a costume for 'technical'" (`:38`), and the hero-metric template (`:26`). Then there is a `<codex>` block that only Codex sees. It bans "sketch-style SVG scenes" and `feTurbulence` grain, `repeating-linear-gradient` stripes, and dismissing things as "theater" (`:44-50`). Those are habits one model family has, written down for that family.

The third pile is the detector, and it is where "vocabulary" becomes something you can run.

<Figure
  src="https://ai.thesatyajit.com/articles/impeccable-design-skill/fig1.png"
  alt="Three before-and-after pairs. Top: a beige card with a purple left stripe, a tracked uppercase label, and an italic serif headline, labelled AI kicker, Italic serif, Side-tab border and AI beige, beside a plain white card with the same text. Middle: a dark analytics panel of four equal metric tiles and a nested weekly chart, labelled Status-chip soup, Everything equal and Cards in cards, beside a white card with one large revenue figure and a bar chart. Bottom: a form headed Unlock your potential today with four fields and a Continue button, labelled Vague headline, Too many fields and Generic CTA, beside a card asking What should this project improve? with three choices and a Use this focus button."
  caption="The ten tells on the home page, in three before/after pairs (/polish, /distill and /clarify), each layer screenshotted on its own in my headless Chromium (impeccable.style home page)."
/>

The ten tells on that hero are a useful test of the three piles. Five of them have a detector rule with the same meaning: side-tab border (`side-tab`), AI kicker (`kicker-above-heading`), italic serif (`italic-serif-display`), AI beige (`cream-palette`) and cards in cards (`nested-cards`). The other five do not. Status-chip soup, everything equal, vague headline, too many fields and generic CTA are judgments about content and hierarchy. Only the model's review in `critique`, or a rewrite in `clarify` and `distill`, can catch them. Half the marketing is measurable, and the other half is the part the model still has to judge.

## The catalogue, counted

The registry is `crates/foundation/src/registry.rs`, mirrored as JSON in `crates/live/assets/antipatterns.json`. It lists 59 rules. Of these, 31 are categorised "slop" and 28 "quality". Their severities are 44 warnings, 13 advisories (listed but never counted in the exit code) and two errors (`script-error`, `content-hidden-at-rest`). The README says 59. The live home page says "61 checks". The `/slop` page renders 67 entries: the 59 rules, six marked "design review" that no code checks, and two the engine has since retired. The site lives in a private repository and trails the engine by a few days.

The "slop" and "quality" labels are the authors' categories. They do not tell you what a rule measures, so I sorted all 65 live entries by that instead. The split is my call, and the widget shows each rule's quoted text so you can disagree with it.

<RuleLedger />

Thirty-three entries are taste written down as a threshold. The authors have judged a habit to be generated, and given it a number so code can find it: cream backgrounds, Inter and Geist, purple gradients, kickers, em-dash density, "theater". Thirteen are readability checks against standards that predate generated UI by decades: WCAG contrast, an 11px floor for functional text, line length, leading, skipped heading levels. Nine are plain defects: a broken image, a script error, text hidden at rest, text clipped or covered. Four compare your values against your own `DESIGN.md`. Six are left to the model.

That answers the question I came with. The vocabulary is mostly a list of refusals. About half of it has been turned into code that measures, and only about a third of the measuring code checks something a non-designer could verify from first principles.

## How a taste becomes a number

The taste rules are worth reading one at a time. Each one forces a vague complaint into a decision someone can argue with. "AI beige" becomes this:

```rust
// crates/core/src/checks/measures.rs:392
pub fn is_cream_color(rgb: Option<&Rgba>) -> bool {
    let Some(c) = rgb else { return false };
    let (r, g, b) = (c.r, c.g, c.b);
    if math_min3(r, g, b) < 209.0 { return false; }
    if !(r >= g && g >= b) { return false; }
    let warmth = r - b;
    warmth >= 6.0 && warmth <= 48.0
}
```

A colour counts as cream when every channel is at 209 or above, red is at least green and green at least blue, and red minus blue lands between 6 and 48. An ivory of `#fbf8f1` has a warmth of 10 and is cream. A warm white of `#f7f7f5` has a warmth of 2 and is not. Nobody would arrive at those bounds from colour science. They are someone's eye, made precise enough to argue about, and that precision is what makes the rule useful.

The side stripe is a geometry rule (`crates/core/src/checks/rules.rs:301`). A coloured border fires on a rounded card at 2px or wider, and on a square one at 3px. The other sides have to be 1px or less, or at most half the stripe's width. Neutral greys are exempt, and so are tables, badges and anything in a status context. The em-dash rule is "advisory only" and needs both at least 8 dashes and at least one per 500 characters of body text (`crates/foundation/src/rules/text.rs:398-401`). The font rule is a list of 17 names, from Inter and Roboto to "the Anthropic-skill / Vercel / GitHub default wave" of Fraunces, Geist and Space Grotesk (`crates/foundation/src/constants.rs:85`). In the browser engine it fires only when one of them sets at least 5% of the page's characters, a floor the code comment says sits clear of the two smallest real cases the corpus found (`crates/core/src/browser/page_checks.rs:112-115`).

I rebuilt thirteen of these patterns on one sample card. Every switch changes what the card renders, and the checker on the right runs a small port of the named rule with the same threshold. The background swatches include that ivory, so you can watch `cream-palette` fire on a colour most people would call white.

<SlopBench />

Two things stood out while I built it. The readability rules have outside anchors: 4.5:1 is the WCAG figure, and nobody has to take Impeccable's word for it. The taste rules have only the authors' eye behind them. Some sit on a knife edge. A beige of `#f3ead8` has a warmth of 27 and would still fire at 47. Move it one notch warmer and it stops being "cream" without looking any different. That is the price of turning taste into a number. The code becomes repeatable, but its edges are arbitrary.

## The clever part is the corpus

What keeps the taste rules honest is not the rules but `tests/oracle/DELTAS.md`. It is 7,015 lines and 109 dated entries, each recording one decision about a rule's behaviour. Many of the decisions come from running the detector on real websites and having two judges label whether each finding actually hurt the page.

> "A corpus review of 50 real sites judged all three rules on their findings. Two measure exactly what they claim and are almost never harmful where they fire (`layout-transition` 1,435 findings on 37 sites, harmful in 9% of the judged representatives; `bounce-easing` 160 findings on 13 sites, harmful in none), so their registry severity is now `advisory`" (`tests/oracle/DELTAS.md:304-308`)

The same entry retires `image-hover-transform` outright, because "hover zoom on a card image is a long-standing convention rather than a generated-UI tell, and it fired on mobile captures where hover cannot happen" (`:309-312`). On 6 October, after two more cohorts in which "both judges called every finding harmless", `layout-transition` was removed from the registry too (`:7006-7008`). The count dropped to 59.

I think this is the best idea in the project. Most lint rule sets only grow. This one gets taken back to the web, measured for harm, and trimmed. Most entries are about precision, not recall. Known consent banners are hidden before a URL scan (`DELTAS.md:4576`). Ad-tech script errors become advisory and name the vendor, sliders that never started are left out of the hidden-text count, and a purple heading that matches the brand's own logo colour stops counting as an AI palette (`docs/CLI-CONTRACT.md:354-359`). Each carve-out is a false positive somebody hit on a real page.

It also shows the cost. Those carve-outs are why the engine is 211,930 lines of Rust across 384 files. The browser rules alone run to 9,004 lines in `element_checks.rs` and 6,891 in `page_checks.rs`. The engine is a port of an earlier JavaScript engine. `docs/PORTING-GUIDE.md:20` says "Port the behavior, including bugs", and 1,410 test files in `tests/oracle/` pin the Rust output byte for byte to the JS goldens.

<Figure
  src="https://ai.thesatyajit.com/articles/impeccable-design-skill/fig2.png"
  alt="A grid of eight small mock interface fragments on pastel backgrounds, each outlined in yellow with a yellow label: AI beige + serif on a large serif headline, Soft rounded card, Side-tab border on an insight card, Eyebrow chip on an INTRODUCING pill above a headline, Pulsing status dot beside 'AI is thinking', Cards in cards on three nested boxes, Identical icon tiles on three feature cards, and Numbered section labels on 01 Discover, 02 Design, 03 Deliver."
  caption="The home page's detector demo: eight tells drawn on purpose, each outlined with the overlay the browser extension draws. Above it the site says '61 checks'; the registry at HEAD has 59 (impeccable.style home page)."
/>

## Three engines, and an ordering rule

`impeccable detect` picks an engine from its target. Source files (`.tsx`, `.vue`, `.css` and so on) go through regex matchers. HTML files go through a static engine with its own parser and CSS cascade. A URL goes to a real Chrome over CDP, which renders the page, sweeps it for scroll reveals, reads computed styles and the painted geometry, and samples screenshot pixels where text sits on an image. The same core compiles to WebAssembly for the Chrome extension, a 2.7 MB generated bundle in `crates/live/assets/`. No engine calls a model, and none needs an API key.

The hooks split the rules into two tiers. After each edit, only "mechanical, unambiguous problems worth interrupting an edit for" are surfaced: broken images, overflow, contrast, gradient text, glow, design-system drift. Copy cadence, palette taste and layout rhythm wait for a deep pass on `Stop` (`skill/reference/hooks.md:7`). This is a good call. An agent that gets told about em-dashes on every keystroke learns to ignore the hook.

The cleverest rule in the skill text governs the model and the detector together. `critique` runs two assessments as separate sub-agents: a design review by the model, and the detector plus a browser pass. The order is enforced:

> "Assessment A must finish before detector findings enter the parent synthesis context. Detector output is deterministic, but it still anchors judgment." (`skill/reference/critique.md:10`)

A run without sub-agents has to start its report with a `⚠️ DEGRADED` banner (`:9`). Whoever wrote that has watched a model read a lint report and then "find" exactly those issues and nothing else. It is the same discipline as checking an agent's work with something that is not the agent, which is what I looked for in [three other harnesses](/articles/agent-tools-week).

## Live mode, the gallery and the dice

`/impeccable live` is the one part that needs a running dev server. The engine injects a picker into your local app (Vite, Next.js, SvelteKit, Astro, Nuxt, Bun or static HTML; no native apps, no deployed sites). You select an element and describe the change. The agent writes three variants that are hot-swapped in place through HMR, and when you accept one, it is written back into your source. Under the hood it is a long-poll loop. `live.md` tells each harness how to hold it: a background task in Claude Code, a one-shot background terminal in Cursor, a foreground exec session that Codex must "SERVICE" (`skill/reference/live.md:26-30`). `generate` does the same for a named element, without the picking.

<Figure
  src="https://ai.thesatyajit.com/articles/impeccable-design-skill/fig5.png"
  alt="Three panels titled Pick, Generate and Accept. Pick: point at what you want to change, with a selected Newsletter card. Generate: explore three alternatives, with a segmented strip of NO.04, DISPATCH and FIELD. Accept: keep the result in source, with a toast reading Variant 2 written to source."
  caption="Live Mode's three steps, as the site draws them (impeccable.style/live-mode)."
/>

The gallery is the closest thing to a benchmark the project publishes. Seven briefs, each run on the same model with and without Impeccable. The page says "first-pass results, before manual polish", with "image generation available to both". Without Impeccable there are five runs per brief, and with it three.

<Figure
  src="https://ai.thesatyajit.com/articles/impeccable-design-skill/fig3.jpg"
  alt="The Observability brief from the gallery. A long prompt for Tortuga, a distributed-tracing tool. Below it, five screenshots labelled Without Impeccable, all cream pages with a black headline about the smallest trace explaining p99 latency. Then three labelled With Impeccable: a transit-map page, a dark green terminal-style page with large type reading P99 WENT UP 212 MS, and a blue-and-white page with a trace diagram."
  caption="The Observability brief with Opus 5.5 selected: five runs without Impeccable converge on one cream page and one headline, while three runs with it produce three different worlds (impeccable.style/gallery)."
/>

The Observability row shows the claim better than any testimonial. The five baseline runs are the same page five times: cream ground, black headline, "the smallest trace that explains" in the first line. The three Impeccable runs differ from each other. That difference comes less from the anti-pattern list than from a mechanism the research page calls "The model can't roll its own dice". Across sixteen creative framings, "Thirty of thirty-five responses proposed the same concept". When the model ranked its own shortlist, "option one won in 27 of 30 packets". So the skill has the model propose directions and has a script pick one: `concept-seed` "assigns the direction to build and deals catalog challengers" (`skill/reference/new-work.md:48`).

The catalog of directions it deals from is not in the repository. It is served from `/api/roll` on impeccable.style. The request carries a scope, a mode, an eight-hex seed key and a re-roll counter, and a separate choice ping honours `DO_NOT_TRACK` (`CLAUDE.md:95`). The repo's own guidance is blunt about why: "Never add catalog data files back to this repo; the catalog is the paid-service moat" (`CLAUDE.md:98`). The research page puts the campaign at "about $2,600 in evaluations", reports a catalog of 375 reviewed entries (188 approved), and says the final workflow "won 100% of decisive pairs" against the strongest competing skill. A separate model judged those pairs, and I could not reproduce any of it, since the evaluation set is private too. Without network access, the roll degrades to the model's own list with the assignment still applied.

## What happened when I ran it on this page

The site owner asked for this article to be built with Impeccable itself: first its widgets, then the whole page. So I treated the cloned skill files as my instructions, ran its passes in the order the references hand off (`critique` and `audit` first, then `distill`, `clarify` and `polish`), and ran its detector before and after. The detector was the published `@impeccable/cli-linux-x64` 0.1.11 binary, downloaded from npm with its integrity hash checked and no install scripts run. I rendered the page in my own headless Chromium and pointed the binary's URL engine at that file, without a model or an API key. The design review in `critique` needs a model. It asks for an isolated sub-agent, so I gave a separate agent the screenshots and source, kept the detector output away from it, and merged the two reports afterwards, as `critique.md:10` requires.

The first draft was what I write by default for this site. It had a bordered panel with a small uppercase "INTERACTIVE · SLOP BENCH" label, a row of three big-number tiles, grey panels inside the panel, and findings as cards with a 3px coloured left border.

<Figure
  src="https://ai.thesatyajit.com/articles/impeccable-design-skill/fig6.png"
  alt="Two versions of the same interactive widget side by side. Left, labelled Before: a bordered panel with a small uppercase INTERACTIVE · SLOP BENCH label, three stat tiles reading 8 findings, 2.5 contrast and 1.15x type ratio, a grey panel of pill toggles and sliders beside a beige sample card, and findings as grey cards with an orange left stripe. Right, labelled After: a plain header with Remove every tell and Add every tell buttons, a sample card wearing yellow annotation tags such as side-tab, cream-palette, kicker and low contrast, sliders whose labels state their thresholds, and a findings list separated by hairlines with collapsed rule text."
  caption="This article's slop bench before and after the Impeccable passes, rendered by me in headless Chromium at 1280px. The left is my first draft; the right is the shipped component with every tell switched on."
/>

<RunLog title="The passes, in the order I ran them" note="Each entry names the command, when it ran and the reference line it acted on.">

<Pass cmd="critique" when="first draft" cite="critique.md:10">

The design review was blunt. It scored the draft 16 of 32 on the heuristics it could apply and called it generic. Its top finding (P0) was that seven of the nine switches changed nothing on the sample card and only added a line to the list: "'Watch the checker react' becomes 'trust me'". Its P1s were that the ledger silently showed 8 of 65 rows, that the draft committed five of the patterns it was teaching (eyebrow, side stripes, nested cards, hero-metric tiles, a ghost shadow), and that the two-column grid had no phone layout.

</Pass>

<Pass cmd="detect" when="first draft">

The detector, run separately, found 46 problems in the rendered page. They included 17 `side-tab` hits, all but one on my own finding cards, and 19 `low-contrast` hits. Sixteen of those were the site's `--muted-foreground` on `--muted`, at 4.3:1.

</Pass>

<Pass cmd="audit" cite="audit.md:58-62">

`audit` told me to run the detector and verify each finding in context. Two were mine: the ledger's rule text ran 81 to 83 characters a line, so it is now capped at 60ch, and the ledger rows put their padding on `summary`, not on the bordered `li`.

</Pass>

<Pass cmd="distill" cite="distill.md:52,56">

`distill` asked for "One font family, 3-4 sizes maximum" and "one spacing scale". The widget's chrome went from six text sizes to four and from eleven spacing values to six, and each rule's full text now sits behind a "Rule text" disclosure.

</Pass>

<Pass cmd="clarify" cite="clarify.md:35">

`clarify` asked that labels "describe what will happen". "Clean card" became "Remove every tell", a slider that read "Grey 150" now reads "Body text contrast 2.5:1 (needs 4.5:1)", and every slider states its threshold.

</Pass>

<Pass cmd="polish" cite="polish.md:73-74,89">

`polish` asked for focus, hover and touch states and a check at every viewport. That pass gave every control a visible focus ring and 44px targets on touch screens. It also caught one problem outside my code: at phone width the site's own figure "expand" button is 10px text at 70% opacity on a blurred chip, which the detector measured at 1.9:1. This page overrides it.

</Pass>

<Pass cmd="detect" when="after the passes">

After the passes the detector reported two findings on the rendered page in light mode, dark mode and at 390px. Neither is a defect. One is `layout-transition` on a site-wide Tailwind utility, a rule HEAD retired on 6 October that the 0.1.11 binary still ships. The other is `theater-slop-phrase`, which fires because the ledger quotes the rule called "Theater framing copy".

With every tell switched on, the detector found ten more, all on the sample card. They overlap the bench's own checker and add two contrast failures I had not built on purpose: the kicker's grey on beige at 4.46:1 and white on the purple button at 3.7:1. The detector missed some of what the bench builds, though, and the misses are instructive. `kicker-above-heading` and `italic-serif-display` look for an `h1` to `h4` or `role="heading"`, and my sample card's title is a paragraph, so neither fired. `cream-palette` only reads the page background. Pointed at the component source, the regex engine found one pattern of thirteen, the `background-clip: text` gradient; a coloured stripe set in a React style object is invisible to it. I added a file-level waiver, `impeccable-disable …: documentation of bad design`, the narrow exception `hooks.md` allows for "documentation of bad design", and the source scan exits 0.

</Pass>

<Pass cmd="typeset, layout, colorize" when="the whole page" cite="mode-read.md:7">

At that point the widgets were done but the page was not. It had the same title block and plain column as every other article on this site, with a wrapper that only narrowed the measure and lifted the contrast. The owner asked for the page itself to be designed, so I worked through the three references that shape a page rather than a component, against the brief `mode-read.md:7` gives a reading surface: "The world owns the frame and serves the column." This time I did the design assessment myself rather than in a separate agent.

The frame is set like a proof sheet. The title puts the product's name at display size and the claim under it, and marks the verdict in the detector's annotation yellow. That yellow is the page's one accent, and it appears elsewhere only on findings and on the command names in this log (`colorize.md:42`). No heading has a label above it (`craft-floor.md:27`). A section opens with a 2px ink rule, a heading hung into the left margin and a first line set in semibold, with four times as much space above the heading as below it (`craft-floor.md:11`). The skill's own sentences are set as pull quotes in the reading face, ruled above and below instead of striped down the side (`craft-floor.md:35`), with the file and line under them. Figures take the column, both margins or the full width, depending on how much detail they carry.

The column stays calm. Body text is Newsreader at 19px, a serif loaded on this page and no other. `typeset.md:57` allows a second family only for a role the first cannot do, and long reading is that role here; the site's Hanken Grotesk keeps the headings and every control. The column is 34rem wide. On a render I counted 69 characters a line at the median, and 63 to 73 on nine lines in ten, against the 60 to 75 that `mode-read.md:15` asks for. Ink and paper are tinted blue in both themes, and I computed every text pair instead of judging it by eye: body text is 17.5:1 and secondary text 7.3:1 in light mode, 16.8:1 and 9.6:1 in dark. Spacing comes from one 4px scale (`layout.md:49`). The only animation is the title's mark being laid down once, in the light theme, and it does not run for readers who ask for reduced motion (`craft-floor.md:13`).

</Pass>

<Pass cmd="detect" when="the whole page">

I ran the same binary on a static render of the new page: its own components and MDX compiled the way the build compiles them, with the site's stylesheet, in light mode, dark mode and at 390px. The site header and footer were left out. The first pass found 41 problems in light mode at 1280px. Twelve were my own. Every section rule was 3px, and a square border on one side counts as `side-tab` from 3px, so the rules are now 2px; and the command names in this log had 2.5px of padding above and below their text. Most of the rest came from the site's shared pieces, which this page now sets in its own terms. The repo card's labels were 9px and 10px, and its tinted header strip read as a card inside the card. The related-article summaries ran to about 124 characters a line.

After those fixes each pass reports four findings, one of them advisory. Two are the known ones, `layout-transition` and the quoted "Theater". One is white on `#ff6600`, the Hacker News mark in the site's share buttons, and I left a brand mark alone. The last is `tight-leading`, at 1.15 on the title's second line. The rule asks for 1.5 to 1.7 "for body text", and this is a display line of up to 44px, so I kept it. That is the judgement `audit.md` asks for: check each finding in context instead of clearing the list.

</Pass>

</RunLog>

Two rules contradict each other, and this page had to pick. The README says "Don't use pure black/gray (always tint)" (`README.md:89`), while `colorize.md:44` says "Neutral gray is valid when it serves the world". The rest of this site is a deliberate pure-neutral ramp and stays one. This page tints, because a blue-black ink is what lets the yellow read as a mark. `craft-floor.md:15` asks for themed custom scrollbars, while `operate.md:50` lists "custom scrollbars" among reinvented affordances to avoid. I themed the selection, focus rings, caret and link underlines, and left scrollbars alone. The sample card keeps literal hex colours in both themes, because the rules read literal colours. Its caption says so.

## How this site compares

This site is also built by agents, and it has its own design step. `brand-crew/agents/design-reviewer.md` is a read-only reviewer. It checks theme tokens, fonts, the shared shell, reduced motion and dark mode, and it is good at keeping pages consistent with each other. It has no opinion on hierarchy, measure or contrast, and that is how 98-character lines and 4.3:1 secondary text had lived in the default styles without anyone noticing. Impeccable's own repository has the prose half of what our `voice.md` does. Its `docs/STYLE.md` is an editorial brief, and its build's `validateProse` step fails on a denylist ("load-bearing", "highest-leverage", "biggest unlock"), the same idea as this site's banned-phrase list. The difference is the visual detector. We lint words; they lint pixels too. The closest thing we have written about is [json-render with Jev](/articles/generative-ui-by-decision), which avoids bad generated UI by never letting the model emit the component tree at all.

## What I would take, and what is thin

The detector is the part I would install. `npx impeccable detect http://localhost:3000` runs on a CPU, needs no model, exits 2 when it finds something, and its readability and defect rules are worth having in CI on any front end. The corpus discipline behind it, where rules are measured for harm on real sites and retired when they are harmless, is what I would copy into my own lint rules.

The skill text is thinner than its size suggests. Much of it is excellent operational detail about harnesses, polling and drift. The design guidance itself is a short list of refusals plus four modes, and 60,000 words of instructions is a lot of context to spend for that. The taste rules will age. The registry already names model families, with `gpt-thin-border-wide-shadow` and `codex-grid-background`, and the font list has a section headed "Newer monoculture" next to the older one. A list of what generated UI looked like this year is a moving target. To their credit, the authors say so on the slop page: "AI slop changes with the era and the model."

The half of the vocabulary that matters most, whether the headline says something and whether everything carries the same weight, is still the model's judgment. The best thing Impeccable does there is procedural. It makes the model judge before it reads the lint report, and it makes a script pick the design direction so the model's favourite does not always win.

## How I checked

I shallow-cloned `pbakaus/impeccable` (`git clone --depth 1`). The clone holds one commit, `d98b0be`, dated 7 October 2026 10:01 -0700, so I could not read the project's history. The npm registry dates the `impeccable` package to 30 March 2026, with 39 published versions; the latest is 4.1.0. I read `SKILL.src.md`, every file in `skill/reference/` and `skill/agents/`, the docs, the build scripts, the launcher, the hooks, the registry and the rule code quoted above, and the `DELTAS.md` entries cited. Rule counts come from `antipatterns.json` and `registry.rs`, severities from the registry, line counts from `git ls-files` piped through `wc`. The catalogue split in the ledger is my own classification. I read every page of impeccable.style (home, Designing, Gallery, Docs, Slop, Live, Research) in my own headless Chromium and took the figures from those renders. The 76k star count is the site's badge; I could not check it against GitHub's API. The only third-party code I ran was the published detector binary (`@impeccable/cli-linux-x64` 0.1.11, sha512 checked against the npm registry, no install scripts), under `nice`, on renders of this article's own components and of this page. The gallery and research results are the project's own, judged by its own reviewers, and I could not reproduce them.
