# Claude's 104 product demos: two clocks, a phoneme-timed mouth and one ffmpeg join

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/claude-demo-videos
> date: 2026-10-06
> tags: agents, tts, video, browser-automation, computer-use, audio

On 6 October Mohd Danish ([@mddanishyusuf](https://x.com/mddanishyusuf/status/2107333673821323412)) posted that he had made 104 demo videos in 40 minutes "and I didn't record a single one. Claude did." His Chrome extension, SuperDev Pro, has 52 developer tools, and each one needed a "how to use" video in landscape and again in vertical for TikTok and Reels. He gave the job to Claude Opus 5.5. It opened a real Chrome on his Mac, installed the extension locally, used every tool "with smooth cursor movement, clicks and typing", logged each step with timestamps, wrote a script per demo, voiced it and assembled the videos. The stack he lists is four lines long: Playwright drives the browser and records, HTML draws every frame, ffmpeg joins video and audio, and Kokoro speaks, locally and with no API key. The background music is generated in code.

I read that list with a slightly uncomfortable feeling, because this site runs almost the same machine. Every article here gets a storyboard, and some get a narrated explainer film: Kokoro voices each line on a CPU, headless Chromium paints the frames, a Python script synthesises the score, and ffmpeg muxes the result. One line of his post did not match ours:

> Kokoro gives the timing of every spoken word, so the mascot's mouth lip-syncs to the voiceover automatically.

Our hosts do not lip-sync. They flap. So I went to find out what Kokoro actually hands you, how much of his pipeline can be reconstructed from what he published, and where ours should change. No code was released (I checked the 20 replies on the first page and the product site), so everything below about his build comes from his post, his own replies and the two example videos he attached. Everything about ours is quoted from `brand-crew/skills/explainer-films/` with line numbers.

## What the frames give away

Two demos are attached to the thread, "Export Element" (24.6 s) and "Extract Images" (25.9 s). Both are 1920x1080 H.264 at 30 fps in the copies X serves, and both follow one template.

<Figure
  src="https://ai.thesatyajit.com/articles/claude-demo-videos/fig1.jpg"
  alt="A white title card. Top left reads HOW TO USE · EXPORT ELEMENT; top right reads SUPERDEV PRO and a running timecode. On the left, a black app icon with a green upload arrow, then 'How to use' in black and 'Export Element.' in green with a hand-drawn underline, and 'Any element → CodePen' in grey. On the right, a large dithered black-and-white mascot with curly hair and round glasses, waving."
  caption="The opening card of the Export Element demo, about two seconds in. The mascot is a dithered bitmap character; the title types itself on (Mohd Danish's demo video, attached to the announcement post)."
/>

I ran ffmpeg's scene detector over both. The title card holds for 4.3 s in the first and 4.6 s in the second. Then comes a three-step section of 13.7 and 14.5 s, and then an identical end card ("Get it free on the Chrome Web Store") that runs 6.6 and 6.8 s. So only the middle half of each video is specific to a tool, which is what makes 104 of them a batch job rather than 104 edits.

<Figure
  src="https://ai.thesatyajit.com/articles/claude-demo-videos/fig2.jpg"
  alt="The Extract Images demo at step 2 of 3. A browser window with traffic-light dots and an address pill reading dub.co shows the Dub homepage, with the SuperDev Pro Extract Images panel open on the right listing 46 images in a grid with a Download all (ZIP) button. A small black cursor sits near the bottom left. On the right of the frame: STEP 2 / 3, the words 'Every image.' in large type, a three-item checklist (Press ⌘K ticked, Every image. current, Download all. pending), a progress bar, and the mascot in a round green-ringed avatar."
  caption="Step two of the Extract Images demo. The browser chrome is drawn, not captured: the page inside it is the recording, and the checklist, headline, progress bar and avatar are an HTML layer around it (Mohd Danish's demo video, attached to the announcement post)."
/>

The step frames tell you more. The window around the page is a drawing: three dots and a rounded address pill reading `dub.co`, with no tabs, no extension icons and no toolbar. That fits Playwright's video, which captures the page's viewport and none of the browser's own UI, so somebody had to draw the window back. The step title, the checklist that ticks itself off, the progress bar and the avatar in the corner are the "HTML draws every frame" part. The top-right corner carries a timecode in hours, minutes, seconds and frames, counting `00` to `29`, so the layer knows which output frame it is drawing.

The extension's own claim checks out. The SuperDev Pro site lists ten featured tools and then "40+ more" by name; I counted 42 names, so 52 in all, as the post says.

<Figure
  src="https://ai.thesatyajit.com/articles/claude-demo-videos/fig3.jpg"
  alt="The closing card: the dithered mascot's head above the SuperDev Pro logo, the line 'Get it free on the Chrome Web Store.' with 'free' in green, a black pill button reading 'Add to Chrome — it's free', and superdevpro.com in small type."
  caption="The shared end card, the last 6.6 s of every demo (Mohd Danish's demo video, attached to the announcement post)."
/>

One detail surprised me. Several frames are two frames at once. You can see it in the corner timecode: in a run of ten consecutive frames, some show a clean value, some show two adjacent values printed over each other, and some values appear twice while their neighbours only ever appear blended.

<Figure
  src="https://ai.thesatyajit.com/articles/claude-demo-videos/fig5.jpg"
  alt="A grid of sixty small crops of the video's corner timecode, ten per row, from 14 to 16 seconds. Many crops show the last two digits doubled, like 02 drawn over 01 or 04 over 03, while others are clean; some values such as 05 and 08 appear in two consecutive crops."
  caption="The corner timecode across sixty consecutive frames from 14 s in, my crops of the published file. Doubled digits are frames blended from two neighbours. In the mid-step frame of the same video (not shown) the whole page is doubled the same way while it scrolls (Mohd Danish's demo video, attached to the announcement post; crops mine)."
/>

A frame-rate conversion that blends neighbours leaves exactly this mark, and the obvious candidate is the gap between Playwright's recorder, which writes 25 fps, and the 30 fps of the published file. I can't tell from the output where the conversion happened, in his pipeline or in X's transcode, but the next section explains why a pipeline built this way is exposed to it. The audio in the file X serves is AAC at 96 kHz. Kokoro speaks at 24 kHz, so the voice was resampled four times over at some point, which costs bits and adds nothing you can hear.

## Two clocks

A demo pipeline has two clocks and they do not agree. The browser runs in real time: a panel takes as long to open as it takes, a page loads when the network says so. Everything else (the overlay, the captions, the voice, the mouth, the music) can be computed from a timestamp and nothing else.

Playwright's recorder lives on the first clock. In the current release, v1.63.0, `packages/playwright-core/src/server/videoRecorder.ts` fixes `const fps = 25;` at line 35, and line 168 spawns ffmpeg with

```text
-f matroska -i pipe:0 -y -an -r 25 -c:v vp8 -qmin 0 -qmax 50 -crf 8
-deadline realtime -speed 8 -b:v 1M -threads 1
```

Each frame is a JPEG from Chromium's screencast, stamped with `frame.frameSwapWallTime` (line 54), the wall-clock moment the compositor swapped. Frames only arrive when the page repaints, so the input has a variable rate, and `-r 25` makes ffmpeg duplicate frames to fill the gaps. A fixed 1 Mbit/s of VP8 is generous for an 800x450 window and thin for 1080p. Main has moved on in the last week: commit `2c751825` on 1 October switched to VP9 and scales the bitrate with the pixel rate ("at 1920x1080 and 60fps the 1M budget visibly blurs scrolling text"), and main now takes an `fps` option too. Danish posted five days later, so his recordings were almost certainly the VP8 kind.

None of that is a criticism of Playwright. A screen recording is a record of what happened, and what happened was not on a frame grid.

Our film engine lives entirely on the second clock. `render.mjs` sets `const FPS = 24` (line 27), and every frame is requested by time: `FILM.frame(t, { blur })` with `t: at[j] / FPS` (line 151). Its header states the consequence, "Frames are pure functions of t, so a film re-renders identically" (line 15). When a film shows a clip, the clip is not played inside the page. ffmpeg pulls it out as JPEG frames at 24 fps beforehand, "so a clip frame is a pure function of time like every other" (lines 52-76).

That last trick is the one I would carry over. The design that keeps a demo deterministic is to cross from the first clock to the second exactly once:

1. Record each tool live, once, with Playwright, and log every action with a timestamp.
2. Explode the recording into frames at the output rate with ffmpeg, and decide there and then how to map 25 fps onto 30: duplicate (sharp, slightly uneven motion) or blend (smooth, with ghosts like the ones above).
3. Paint every output frame as a function of `t`: the recorded frame for `t`, the overlay, the cursor, the captions and the mouth.

After step 1 nothing depends on wall-clock time, so the landscape and vertical versions are two layouts of the same frames and the same audio, and a rerun gives the same file. One more thing to get right in step 1: Playwright's video starts at the recorder's own creation time (`this._creationTimeMs = Date.now()` in the constructor, line 117), and your log starts whenever you call `Date.now()` in your script. Those are a few hundred milliseconds apart and the gap is not fixed. Put a visible marker on the page at the moment your log says zero, such as a one-frame colour flash, and find it in the video, rather than trusting two clocks to agree.

The time budget is the other reason I would split it this way. Recording the 104 finished videos in real time would by itself take 104 × 25 s, about 43 minutes, more than the 40 he quotes. Recording 52 tools once and rendering both layouts from the same footage fits comfortably: 52 step sections of about 14 s each is around 12 minutes of browser time.

## The cursor nobody recorded

The pointer is not part of the page, so Chromium's screencast never contains it. In Figure 2 the small black arrow in the lower left was drawn by the overlay. That leaves you free to draw a better cursor than the one Playwright moves.

Playwright's own motion is plain linear interpolation. In `input.ts` (v1.63.0, lines 222-230):

```ts
const { steps = 1 } = options;
// ...
for (let i = 1; i <= steps; i++) {
  const middleX = fromX + (x - fromX) * (i / steps);
  const middleY = fromY + (y - fromY) * (i / steps);
  await this._raw.move(progress, middleX, middleY, /* ... */);
}
```

With the default `steps = 1` the pointer teleports. With more steps it moves in equal increments, sent as fast as the protocol round-trips, so `steps` sets how many events the page receives and not how long the move takes. For hover states and drag handlers that is what you want. For a viewer it looks like a machine.

So drive the page with Playwright, log where each move starts, where it ends and when, and draw the cursor yourself from that log with a human motion profile. [Cua's cursor work](/articles/arc-cua#update-cuas-six-cursor-motions) is a good reference here: its arc styles start from $150 + 120\,\log_2(D/W + 1)$ ms, Fitts' law clamped to 300–1000 ms, where $D$ is the distance and $W$ the target's width, and eases along a curve with a minimum-jerk profile. Minimum jerk is the classic model of a human reach:

$$
s(u) = 10u^3 - 15u^4 + 6u^5,\qquad u \in [0, 1]
$$

Its speed and acceleration are zero at both ends, so the cursor leaves gently, rushes the middle and settles onto the button. The widget samples one dot per 30 fps frame; the gaps between dots are the speed a viewer sees.

<CursorEasing />

Take the far, small target: 520 px to a 28 px button, which Fitts' law gives 665 ms, or 20 frames. The eased cursor moves 0.6 px in its first frame and about 48 px in its fastest; the linear one moves 26 px in every frame and arrives at full speed. The arc matters less than the ease. One practical note: if Playwright's real pointer and your drawn one disagree about where a click happened, the click wins, so draw the drawn cursor's arrival a frame or two before the logged click time, never after.

## What Kokoro already knows about time

Now the part that sent me down this road. Kokoro-82M is a StyleTTS 2 descendant with an explicit duration model, and that model is where the timings come from.

In `kokoro/model.py` (hexgrad/kokoro at `dfb907a`, lines 108-110), after the predictor's LSTM:

```python
duration = torch.sigmoid(duration).sum(axis=-1) / speed
pred_dur = torch.round(duration).clamp(min=1).long().squeeze()
indices = torch.repeat_interleave(torch.arange(input_ids.shape[1], device=self.device), pred_dur)
```

Every input symbol, phonemes and stress marks and spaces alike, gets a whole number of frames. `repeat_interleave` turns those counts into an alignment matrix, and the decoder generates audio from the aligned features. The durations are therefore not an estimate of where the sound lands. They are the plan the audio is built from. The pipeline's comment gives the unit: "Multiply by 600 to go from pred_dur frames to sample_rate 24000" (`kokoro/pipeline.py:296`), so one frame is 25 ms.

I checked that on real output. I ran the 8-bit ONNX export of Kokoro-82M with the `af_heart` voice at speed 1.1 (the speed our films use) on the line "Press command K, then pick an element, and send it straight to CodePen." with the duration tensor exposed as an extra output. The predictor gave 186 frames and the decoder returned 111,600 samples, exactly 600 × 186, or 4.65 s. The other two speeds I tried matched to the sample as well.

Word timestamps are a sum on top of that. `KPipeline.join_timestamps` (`pipeline.py:295-331`) walks the tokens, adds up each word's phoneme frames, splits each space's frames between its neighbours by counting in half-frames, and writes `start_ts` and `end_ts` onto each token. Two details are worth knowing before you build on it.

- The start is shifted. `left = right = 2 * max(0, pred_dur[0].item() - 3)` drops three frames of the leading pad, next to the authors' own `# TODO: Is -3 an appropriate offset?`. On my line every word's `start_ts` sits 75 to 113 ms before the frame where the decoder starts that word. For a mouth a small lead is harmless, since animators often open the mouth a frame early, but if you align captions to the audio and then also lead them, the two leads stack.
- Timestamps only exist for English. The `lang_code in 'ab'` branch runs `join_timestamps` (lines 394-395); every other language takes the chunked branch, which yields a `Result` with no tokens (line 442). The durations are still computed there, just not handed back as tokens, so you would read `result.pred_dur` yourself.

Speed is not a simple stretch either, because of `clamp(min=1)`. At speeds 0.9, 1.1 and 1.3 the same line took 216, 186 and 162 frames. Pure scaling from 186 would give 227 and 157. Many phonemes already sit at one frame, and a phoneme cannot get shorter than one frame, so speeding up compresses the vowels and leaves the consonants alone.

The widget below plays the line and drives two mouths from it. The left one follows Kokoro's frames; the right one is our film engine's mouth.

<NarrationTiming />

## From frames to a mouth

Lip-sync from phonemes is old and simple. You map each phoneme to a mouth shape, a *viseme*, and hold the shape for the phoneme's frames. I used a reduced set of seven: rest, lips shut (p, b, m), lip on teeth (f, v), teeth (s, t, d, n, k and friends), wide (i, A), open (æ, ɑ, ə, ɛ) and round (u, O, ɹ). A stress mark has frames but no shape of its own, so it borrows the shape of the vowel it stresses. Danish's mascot has at least three mouth shapes that I can see in the published frames, plus a blink:

<Figure
  src="https://ai.thesatyajit.com/articles/claude-demo-videos/fig4.jpg"
  alt="Twenty consecutive crops of the dithered mascot's face from the round avatar, two rows of ten, at fifteen per second. The mouth changes between a closed smile, a small open mouth and a wider open mouth showing a tongue; in one frame the eyes are closed in a blink."
  caption="The avatar's face from 5.5 s, fifteen crops a second. The mouth moves between a closed smile, a small opening and a wide one with a tongue, which is a three- or four-shape rig driven per phoneme or per word (Mohd Danish's demo video, attached to the announcement post; crops mine)."
/>

He says word timings drive it. With only word spans you can open on a word and close in the gaps, which reads well enough at avatar size. With the phoneme frames you get closures on p, b and m, which is what makes a mouth look like it is saying the word rather than talking in general.

Now ours. `engine/explainer.js:49-52`:

```js
function talkAt(t) {
  if (!TL) return 0
  for (const b of TL.beats) if (t >= b.t && t < b.t + b.d) { const x = t - b.t; return .3 + .7 * Math.abs(Math.sin(x * 15)) * (.55 + .45 * Math.sin(x * 4.3)) }
  return 0
}
```

While a line plays, the mouth opening is a sine with a slower sine on top, never below 0.3, and `mascot.js:515` draws any value above 0.08 as an open mouth. Across a whole line the host never closes its mouth. The line in the widget has five lip closures (the p in "Press", the m in "command", the p in "pick", the m in "element" and the p in "CodePen") and our host shows none of them.

Captions have the same problem in a different place. `audio.py`'s `captions()` (lines 530-546) cuts a line into cues of about twelve words and times each cue "by word share", splitting the line's duration in proportion to word count. On the widget's line, a 13-word sentence, an even split puts "CodePen" 0.45 s late, about eleven frames at 24 fps, because a one-syllable word gets as much time as a four-syllable one. Per-cue that is tolerable. Per-word highlighting, the kind his demos and every short-form video use, needs the real spans.

The fix in our code is one line wide. `audio.py:106` voices a line as

```python
chunks = [a for _, _, a in pipe(speakable(line), voice=args.voice, speed=SPEED)]
```

That loop unpacks each `Result` through its backward-compatible three-tuple (graphemes, phonemes, audio) and discards the tokens. `join_timestamps` has already run by then, every time, on every line of every film. We compute the timings and throw them away in an underscore. Keeping them means writing `r.tokens` and `r.pred_dur` beside each wav, offsetting each chunk's times by the length of the chunks before it (a long line comes back in several chunks, each timed from its own zero), and subtracting what `trim()` (lines 115-119) cuts from the front of the line. That last one is easy to miss: `trim` removes leading silence down to 30 ms before the first sample above 0.01, so the wav no longer starts where Kokoro's clock does.

The sketch below targets the PyTorch package. I did not run this exact code, because the PyTorch stack is not installed on this machine; the numbers in this article come from the ONNX route described at the end.

```python
from kokoro import KPipeline
import numpy as np

SR, FRAME = 24000, 600 / 24000
pipe = KPipeline(lang_code="a")

def voice_with_timings(line, voice="af_heart", speed=1.1):
    audio, words, phones, t0 = [], [], [], 0.0
    for r in pipe(line, voice=voice, speed=speed):
        for tk in r.tokens or []:
            if tk.start_ts is not None:                  # punctuation-only tokens can be None
                words.append((tk.text, t0 + tk.start_ts, t0 + tk.end_ts))
        dur = r.pred_dur.tolist()                        # [bos, *one per phoneme symbol, eos]
        t = t0 + dur[0] * FRAME
        for ch, d in zip(r.phonemes, dur[1:-1]):
            phones.append((ch, t, t + d * FRAME))
            t += d * FRAME
        a = r.audio.numpy()
        audio.append(a)
        t0 += len(a) / SR                                # the next chunk is timed from its own zero
    return np.concatenate(audio), words, phones
```

## Who writes to whose clock

The deepest difference between his pipeline and ours is the order of work, and it follows from the subject.

Our films are voice-first. `build.mjs` lays it out in its header (lines 22-32): the narration lines come from the storyboard, Kokoro speaks them all, "each beat lasts the longer of its reading time and its speech", and only then are frames painted. The picture waits for the voice. `beatTimes()` in `explainer.js` (lines 41-45) gives each beat its spoken length plus half a second, with a 0.35 s lead before the first. Because the picture is timed to this exact recording, the mixer can refuse anything else. `audio.py` (lines 556-574) starts each line at its beat or just after the previous one ends, and if any line would slip more than 0.35 s it exits with "The film was not timed to this voice". It is there because, in `build.mjs`'s own words, "a film painted without the voice pass is timed to guesses, and its lines talk over each other" (lines 17-19).

A product demo cannot be voice-first, because the browser owns the timeline. The panel opens when it opens. Danish's order, as he tells it, is action first: use the tool, log timestamps, *then* write the script. Each step's narration then has a slot whose length is already fixed, and there are three ways to fit it:

- Write to the slot. Kokoro at speed 1.0 reads "~2.8 words a second" (`audio.py:27`), a little faster at our 1.1, so a 4 s step holds a dozen words at most, and ten is a safer target. Have the model write to a word budget, voice it, measure, and rewrite if it runs over. Our storyboards do the same at film scale: at most 34 words a beat and 280 a film (`SKILL.md`, lines 90-96).
- Stretch the picture. Hold the last recorded frame of a step while the voice finishes, or slow a quiet stretch slightly. Because everything after the recording is a function of `t`, this is a remap of time, not a re-record.
- Never stretch the voice. Kokoro's `speed` is for the whole line, and as the clamp showed, it does not scale evenly.

His reply about redos fits this picture. He made one video first, "and that took 3 redo", then told Claude to make the rest, and the remaining 103 needed none. A template that has survived three iterations of slot-fitting is what lets the batch run unattended. It also means the 40 minutes covers the batch and not the three rounds before it.

## Music with nothing to license

"Even the background music is generated in code," he writes, so there is nothing to license. Ours makes the same choice for the same reason. `audio.py`'s docstring (lines 18-21): "Nothing is sampled or licensed from anywhere: every effect and every note of the score is synthesized below from sine waves, noise and a plucked-string model, seeded by the film's slug, so the same storyboard always sounds the same."

You do not need much. The crayon style's music box (`score_musicbox`, lines 429-438) is a glockenspiel tone (a sine plus two inharmonic partials at 2.76 and 5.4 times the fundamental, each decaying), an eight-step arpeggio over a four-chord loop, and a whisper of band-passed noise on the eighths. Seeding by slug makes the key and the progression differ between films while staying reproducible, which matters when a re-render should change nothing you did not touch.

Mixing is where code music usually goes wrong, and our mixer sets levels by measurement rather than by synthesis gain (lines 587-594): the voice at RMS 0.1 while speaking, the score scaled to RMS 0.032, about 10 dB under, then ducked by up to a further 60% wherever a 0.3 s smoothed voice envelope is active. ffmpeg's `loudnorm=I=-16:TP=-1.5:LRA=11` (line 601) brings the whole film to a common loudness. For 104 demos played back to back on a feed, normalising loudness matters more than which chords you picked.

## The join, and what to encode with

The last ffmpeg call is the easy one. Ours, `build.mjs:130`:

```bash
ffmpeg -y -i frames.mp4 -i mix.m4a -map 0:v -map 1:a -c copy -shortest -movflags +faststart out.mp4
```

Stream copy, no re-encode, the shorter stream sets the length, and `+faststart` moves the index to the front so a browser can start playing before the download ends. The video was already encoded once, from the frames, and that is the encode to get right.

Our encoder settings (`render.mjs:103-114`) are tuned for our content and would be wrong for his. We use x264 with `-preset slow -crf 32 -tune animation`, downscaled to 960x540, because the films are flat drawn art on an article column about 700 px wide. Across the 63 rendered films in our manifest, 86 to 124 s long, the files average 2.32 MB, about 167 kbit/s with the audio. A UI recording is the opposite case: small text, thin borders and scrolling are exactly what a high CRF smears first, and `-tune animation` assumes flat colour areas a web page does not have. For demo footage I would start well below our CRF, use the default tune, and judge it on the frame with the smallest text in it, since that frame fails first.

Two more cautions for this kind of pipeline. Do not let the intermediate be lossy twice: Playwright's VP8 recording is one generation already, so explode it to frames once and encode once at the end. And for social platforms, assume they will re-encode whatever you upload. The 1080p file X serves for his first demo is 7.0 MB for 24.6 s, but that is X's encode, and it says nothing about the file he uploaded.

## Things that bit us

The four problems in our engine's history are the ones a pipeline like his will run into next.

Names first. Kokoro reads through a dictionary, and a word missing from it gets spelled out: "Qwen" came out "Q-wen" (`SKILL.md`, lines 98-106). Product names are the worst case, and a demo for a developer extension is full of them: CodePen, Dub, ZIP, a keyboard shortcut. We keep a table of phoneme overrides in `pronounce.json`, which Kokoro accepts inline (the word in square brackets, then its phonemes between slashes in parentheses), and `audio.py` refuses to voice a film that contains a name the voice would have to guess (lines 89-91). The captions show the text as written, so the fix goes in the dictionary, never in a respelled script.

Budgets second. A film with no word limit drifts long, because a model describing what it sees always has one more thing to say. Ours caps a beat at 34 words and a film at 280; his demos have a harder cap, the length of each recorded step.

Rendering third. Our brush styles paint with WebGL, and on a machine with no GPU Chromium quietly falls back to a software renderer. Which one you get matters: `render.mjs` (lines 179-226) starts a private Xvfb so Chromium can use Mesa llvmpipe, then asks the page which renderer it actually got, because "Without one Chromium quietly falls back, so the renderer is checked, not assumed." The commit that introduced it took a brush bake from about 580 ms to about 52 ms. His overlays are plain HTML and CSS on a Mac with a GPU, so this one may never reach him. On a headless server it would.

Clips last. A two-second clip looped inside a scene "restarting its counters and showing values between the ones the narration names" (commit `1753810`); a clip now plays once and holds its last frame. Recorded UI has the same trap. If a step is shorter than its narration, hold the last frame. Do not loop the footage, or the viewer sees the tool reset mid-sentence.

## If I built it tomorrow

Taking the best of both, I would work in this order:

1. Record each tool once with Playwright at the viewport size of the largest layout, logging every action with a timestamp and putting a sync flash at log zero.
2. Explode each recording to frames at the output rate.
3. Have the model write each step's line to a word budget from the step's logged length, voice it with Kokoro, and keep `tokens` and `pred_dur`. Rewrite any line that runs over its slot.
4. Paint every output frame from `t`: recorded frame, drawn window, eased cursor from the log, captions from word spans, mouth from phoneme frames. Landscape and vertical are two layouts of this one function.
5. Synthesise the score from a seed, duck it under the voice, normalise loudness, and mux with stream copy.
6. Gate before the full batch: render one tool, look at a contact sheet, and only then run the other 51. He did this by hand with three redos; ours does it with `validate:films` and a contact sheet, and [the launch-video article](/articles/one-shot-launch-videos) shows two other ways of gating.

What I would not claim is "40 minutes" without the setup. The batch time is real, and the per-video cost of a good template really does approach zero. But the template, the three redos and the extension install all happened first, and the ratio that matters to anyone copying this is setup time against batch size. At 104 videos it is a very good trade. At five, I would record them by hand.

As for our films: their hosts are going to stop flapping. The timings were already there, one underscore away.

## How I checked

- **The post and thread.** Read through the fxtwitter mirror: the main post, the author's one follow-up (a second demo), and the first page of 20 replies, which includes his answers on redos ("first I ask to make only 1 video and that took 3 redo") and his plan (Max). No code, repository or prompt was shared there, and the product site links none. Everything about his internals beyond the four-line stack in the post is inference from the frames, and I have said so where it is.
- **The demos.** Downloaded both attached videos at 1920x1080 from X's CDN. Durations, frame rate and audio sample rate are from `ffprobe`; section lengths from ffmpeg's scene detector (threshold 0.12); the blended timecode and the mouth shapes from frame crops I made and looked at. These describe the files X serves, which may differ from what he uploaded.
- **The tool count.** Rendered superdevpro.com in a headless Chromium (it sits behind a Cloudflare check) and counted the tool names: 10 featured and 42 listed, 52.
- **Playwright.** Read `videoRecorder.ts` and `input.ts` at the v1.63.0 tag and `videoRecorder.ts` on main (`d469960`), and the file's recent history through the GitHub API.
- **Kokoro.** Read `kokoro/model.py` and `kokoro/pipeline.py` in a shallow clone of hexgrad/kokoro at `dfb907a`. To get real durations without the PyTorch stack, I loaded `onnx-community/Kokoro-82M-v1.0-ONNX`'s `model_quantized.onnx`, added the graph's `/encoder/Clip` tensor (the rounded, clamped durations) as an extra output, phonemised the line with espeak-ng and mapped it to Kokoro's alphabet, and ran it with the `af_heart` voice at speeds 0.9, 1.1 and 1.3. Word spans come from my port of `join_timestamps`. Because the phonemiser and the quantised weights differ from the PyTorch pipeline's misaki and full-precision model, individual frame counts will differ slightly from what `KPipeline` returns; the structure (600 samples a frame, the -3 offset, the one-frame floor) does not. The widget plays those three generated clips.
- **This site's engine.** Every line number refers to `brand-crew/skills/explainer-films/` at commit `856fbee`. The film sizes are summed from the films manifest.

<ModelCard repo="hexgrad/Kokoro-82M" />

<RepoCard repo="hexgrad/kokoro" />
