2026-10-06 · 16 min · explainer · agents · harness · open-source · video · rust
Three repositories went round this week with the same shape. Each takes a coding agent that already
exists (Claude Code, Codex, DeepSeek's dsh) and wraps it around one domain: games, video, the
agent's own launcher. The posts were short:
Genex promised a "self-improving harness for
game dev", ffmpeg-skill promised 42 tools that
probe first and verify after, and rdsh promised
81x faster startup in 799 KB.
I cloned all three with git clone --depth 1 and read them. I did not install or run any of them.
Every number below is labelled: measured (I counted it in the source), reported (the project
or its post says so; I did not re-run it) or reasoned (my arithmetic on the other two).
The thread that ties them together is the one I care about. A harness, in the sense of
Lilian Weng's framing, is the loop and scaffolding around a model. The
interesting part of each of these three is where that scaffolding makes the agent check its own
work, and what kind of check it is: a compiler, a blind judge, an ffprobe, a picture, or a
benchmark.
- license
- MIT
- branch
- dev
- tests
- 672 files
- source
- 29.1 MB
- commit date
- 2026-10-05
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-06 at f23e1c3 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
- license
- MIT
- branch
- main
- tests
- 29 files
- source
- 2.3 MB
- commit date
- 2026-10-05
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-06 at 008333a — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
- license
- MIT
- branch
- main
- tests
- 5 files
- source
- 478.0 kB
- commit date
- 2026-10-06
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-06 at 8b3a8d5 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
Genex: a game studio that edits its own instructions
genex-games/genex-desktop is an Electron app, MIT,
copyright genex.games, read at commit f23e1c3. Its README links the same genex.games site the
post does, so it is theirs. A person describes a game in a chat on the left; a stage on the right
shows the running game, its builds and its assets.

How it plugs into a subscription
There is no model API key in the main path. src/substrate/engines/claude-code.ts drives Claude
Code through @anthropic-ai/claude-agent-sdk (^0.3.257 in package.json, measured). Its header
comment states the boundary in so many words: the studio does not route a consumer subscription
through a model API itself; "the SDK spawns Claude Code, which performs its own login and keeps its
own credentials in its own config directory." codex.ts does the same for ChatGPT by spawning
codex exec --json and reading its JSONL event stream. Both strip ambient keys: an
ANTHROPIC_API_KEY in your shell does not count as configured, and CODEX_API_KEY and
OPENAI_API_KEY are removed from every child environment, so a stray key cannot flip the bill to
per-token (measured from the code and its comments).
Local models go through two more engines: a managed Bonsai 2 27B on Apple Silicon Macs with at
least 16 GiB, and Ollama by exact model name (reported, docs/local-models.md). Meshy, Tripo and
ElevenLabs are not local. They come through a "Genex Tools" plugin whose manifest reads "made in
chat with Genex credits. Coding stays on your own provider." That is the business model, and it is
stated plainly.
How it drives Three.js and Blender
The games are browser games on Three.js (three ^0.185.1). The harness does not trust a game to
describe itself. Every page the studio serves gets a shim before the game's first line, so
window.__studio owns the clock, makes seed(n) reproducible and counts draw calls at the graphics
API (reported, docs/harness-runtime.md). Phaser, plain canvas 2D and engine exports are out of
scope for that contract.
Blender is a plugin, and its manifest is mostly a command line. The model tool takes a bpy
script the agent writes and runs it as blender -b --factory-startup -noaudio --python-exit-code 1 --python wrapper.py -- <script> <model.glb> <render.png> <name>, with a 120,000 ms timeout and a
104,857,600-byte (100 MiB) cap on assets (measured, src/plugins/blender/plugin.json). The
wrapper exports model.glb and two renders, render.png and render-front.png, and the backend
returns those renders to the agent as images. The skill text attached to the tool says what to do
with them: inspect the renders, load the exact GLB path with GLTFLoader, then "verify it in the
actual game preview." If Blender is missing, the plugin offers a pinned Blender 5.2.1 download of
346,264,899 bytes, checked against a SHA-256 in the manifest.
What "self-improving" is in the code
This is the claim worth reading for. The harness that builds games is not compiled into the app. It
ships as a seed (src/harness-seed/): TypeScript that Node runs by stripping its types, copied into
a workspace the in-app agent owns. The docs name four places learning can land, "and they are not
interchangeable" (reported, docs/harness-runtime.md):
| Where | What it holds | In the seed (measured) |
|---|---|---|
library/checks.json | executable checks on a running game | 5 checks |
library/recipes/*.json | craft opinions, each with an optional check | 44 recipes |
skills/*.md | how an agent works | 2 skills |
prompts/*.md | what a role is | 3 files |
The checks are code, not prose. drawcalls-ceiling is the expression
__render.drawCalls <= 1000 && __render.triangles <= 400000; player-moved drives the game's own
movement keys and asks whether player.x or player.z changed. A recipe pairs a paragraph of
intent with a check. characters.feet-on-the-ground says to place a figure by its base rather than
its centre, and checks that every object tagged enemy or npc has its bounding-box bottom within
5 cm below or 12 cm above the ground mesh. Its stats in the seed are applied: 0, wins: 0, status
candidate (measured). A recipe is retrieved when a check fails or a judge names the defect, "never
imposed".
There are two ways those files change.
The agent edits itself. tools/self-tools.ts gives the in-app agent write_own_file,
write_skill and install_tool. None of them writes directly. Each goes through
guardian.write_self: the host makes a validation fork of the harness, writes the change there,
runs a vendored TypeScript 7 tsc --noEmit inside the sandbox with a time limit, boots the fork and
asks its healthcheck. Only a pass is written, between two snapshots, with an Activity entry the
person can undo. A change may not add type errors; errors the code already had do not block it. The
judge rubrics are out of reach: judge/ is write-denied to every agent process, and
write_own_file refuses it by name with "judge/ is frozen: the blind critic's rubric is not yours
to edit" (measured).
A loop edits the skills between runs. loop/skillopt.ts is a port of Microsoft's SkillOpt, and
its header lists what it kept: bounded edits, a ranking, a gate, and a buffer of refused ideas. It
mines up to 12 replayable tasks from the run log (facet iterations, refused gates, finished workers, landings), each marked a success or a failure,
splits them every other one into a train half and a held-out half, and asks an analyst model for at
most 4 edits from four string ops (append, insert_after, replace, delete), never inside a
<!-- SLOW_UPDATE --> region. The candidate skill then faces a blind pairwise gate on the held-out
tasks: 3 votes, the candidate swapping sides between votes so a judge with a position bias cannot
hand it a sweep, accepted only if it wins more than half (measured). A refusal goes into a step
buffer of up to 50 so the analyst is shown its dead ideas next time.
Two honest limits sit in the source itself. The split exists because of a failure: "When the analyst saw every task, its candidate restated the very failures the gate then asked about, and all 55 candidates a real install staged won 3/3" (reported, a code comment). And the gate judges text, not games: the docs say skill gates "compare instruction texts against saved task descriptions; they do not execute candidate builds or establish better future game outcomes." So a SkillOpt pass improves what a blind model thinks of the instructions. Whether the next game is better is a separate measurement, and the product copy says as much: "A learning count does not certify a better game." Learning is on by default; automatic application of suggestions is off on fresh installs (reported).

That is a more careful design than the phrase suggested. The agent is allowed to change its own body, but every change passes a compiler and a boot, the yardstick is frozen, and the outer loop keeps an edit only against tasks it did not learn from. It is the same move as the harness-as-generalizer result, where the scaffold rather than the weights is what learns, and it carries the same risk Skill2Env ran into: a skill judged good by a model is not yet a skill shown to help.
ffmpeg-skill: the agent's hands, and a rule to look
kajisho5/ffmpeg-skill is version 2.5.1, MIT, read at
008333a. It is an Agent Skill:
a SKILL.md and a scripts/ directory of Python 3.9 standard-library scripts that wrap ffmpeg
and ffprobe. npx ffmpeg-skill copies it to ~/.claude/skills/ffmpeg-skill (--cursor,
--codex or --all for the others).
Counting the tools. scripts/ holds 42 Python files whose names do not start with an
underscore, and scripts/_contract.py describes exactly the same 42, no more and no fewer
(measured; I parsed the contract table rather than importing it). By role: 32 execution, 3
analysis, 3 analysis-and-execution (silence, loudness, sync) and 4 verification (check,
look, verify, report). 27 tools are flagged visual, and every one of them lists look among
the checks to run afterwards (measured). Over MCP, tools/list shows a core 12 by default to save
context; the other 30 are still callable by name.

Probe, plan, run, verify, in the code
The post's "probe first, then process, then verify automatically" is real, but it is two layers, and the split matters.
The first layer is code. Every writing script ends in emit(), which calls verify_output() on
the file it just wrote: the file must exist, must not be 0 bytes, and ffprobe must read at least
one video or audio stream from it. Otherwise the script dies with kind: output. The JSON result
then carries "verified": true only when the file was written, probed and every self-check the
tool added passed; a dry run verified nothing. A few tools add their own measurement, such as
loudness after a write.
The second layer is instructions. SKILL.md makes the agent probe before planning, run with
--dry-run --json to see the exact command lines, compare the output's probed duration,
resolution, fps and audio against the request, and, whenever the picture changed, run
look.py OUTPUT --tiles 3x2 and view the PNG. "Not finished until Look: names that PNG — a probe
cannot see a caption on a face." With no vision, the agent must say so:
Look: PATH (pixels not inspected; agent has no image view).
So the code proves the output is media; whether it is the right media is a rule the agent follows. Here is one job through both layers, with the command line each script builds:
Workflow step 1. Every later number comes from this, not from the file name.
python3 $S/probe.py talk.mp4 --compact
ffprobe -v error -print_format json -show_format -show_streams -show_chapters talk.mp4
talk.mp4: 30.000 s, 1920x1080, 30 fps, h264 + aac stereo, VFR suspected: no
The widget's clip is made up: a 30-second talk with five quiet spans. The commands are not.
silence.py asks silencedetect for spans quieter than -35dB for at least 0.6 s, then walks
them with keep_ranges(), which keeps a --margin of 0.15 s of silence either side of speech and
drops kept pieces shorter than --min-keep 0.2 s. The kept ranges become one expression,
between(t,a,b)+between(t,c,d)+…, fed to select for video and aselect for audio, re-encoded in
one pass at CRF 18. Moving the margin slider shows the trade: more margin keeps breaths and
lengthens the result; less margin tightens it and starts clipping word edges.
Two things the stepper surfaces that the post does not. First, order matters and the skill knows
it: silence comes before captions in its chain, and caption.py has an --offset flag but no way to
remap cues through silence.py's cut list, so subs.srt has to be timed against the tightened
file (or transcribed from it with --transcribe and a local whisper). Second, "verify" for the
silence step is structural. The script logs expected ~N s; comparing that against the probe is the
agent's job, not the script's.

The project's own evidence is reported, not re-run: 0 missed gaps on 20 silence.py cases with
known gaps, 92 of 92 verification steps on a 10-file real-device corpus, and 72 of 72 agent runs of
24 prompts graded by an independent model, including the visual check run 24 of 24 times the picture
changed (README, "Tested on real footage"). The evals/results/ directory keeps those runs as JSON,
in iterations numbered up to 25, which is the right habit even if I did not check the grades.
Compared with the launch-video skills, which hand ffmpeg a
finished frame sequence at the end, this is the inverse design: the model never writes a filter
graph at all. "Nothing runs through a shell; no filter graph is accepted from the caller." The skill
also tells the agent not to fall back to raw ffmpeg when no script covers a request, because that
"bypasses every guarantee this skill makes."
rdsh: a fast path that mostly doesn't run the agent
sahenjp/rustdsh builds a binary called rdsh, version
0.1.3, MIT, read at 8b3a8d5: 7,214 lines of Rust under src/ (measured). It launches dsh,
DeepSeek Harness, the plugin-everything agent harness that
deepseek-ai/deepseek-harness ships on npm as
@deepseek-ai/dsh and that needs Node ^22.19.0 || >=24.0.0 (measured, its package.json).
What rdsh replaces is the command line in front of dsh, not dsh. Its architecture note is one
diagram: argv goes either to a native fast path (tokens, search, compact, sessions,
guard, auth, doctor, a status page) or, for anything agentic, to "passthrough: verbatim exec
of dsh-orig". On Unix that is a real exec(3): rdsh tui replaces itself with the original dsh,
and Node boots exactly as it would have. Installed with --as-dsh, the binary shadows the dsh
name and keeps the original as dsh-orig.

What the benchmarks measure
The post's numbers (reported): startup 81x, memory 1/23, search 3.4x, 799 KB against a Node tree of about 508 MB. The README has since moved: about 0.90 ms against 88 ms for startup, which it calls about 98x; 2.9 MB against 66 MB peak memory; search 17 ms against 41 ms, about 2.4x; one binary of about 806 KB (reported, Linux x86_64, median of n=5).
The method is in src/main.rs, and it answers a narrower question than the post implies:
- Startup is
rdsh --versionagainstdsh --version, each spawned 5 times and the median taken. rdsh's--versionis printed byclapand exits; the comment in the source says so ("--version/-V is served by clap itself").dsh --versionboots a Node runtime. So the 98x is a real measurement of how long it takes to print a version string. It is not how long an agent session takes to start, becauserdsh tuistill execsdshand pays Node's boot plus about a millisecond (reasoned). - Search, tokens and sessions compare rdsh with rdsh: "before" is a build of rdsh's previous
HEAD, "after" is the working tree, outputs diffed for equality. The 2.4x is an optimisation inside
rdsh. There is no search-against-
dshrow. - Memory is the same
--versioncomparison under/usr/bin/time -v. - Size compares rdsh's binary with
dsh's Node install, but rdsh needs that Node install to do anything agentic, so the two do not substitute (reasoned).
"Slim mode" also does less than its name. slim_env() sets six RDSH_* variables, and the file's
own comment says upstream dsh reads none of them; I found no RDSH_ string anywhere in the
upstream checkout (measured). The one variable that reaches Node is NODE_COMPILE_CACHE, which
caches compiled V8 code on Node 22.1 and later. That can shorten a real dsh boot. The README does
not measure by how much.
The repository's docs have drifted from its code in small ways: the README says 26 unit tests and
"only three dependencies", while src/ holds 62 #[test] functions and Cargo.toml lists four
crates, getrandom included (measured). None of that is a defect in the binary. It is what a
project that moved from 81x to 98x in a few days looks like.
The useful part for an agent is small and honest. rdsh guard is a PreToolUse hook command:
it scans the hook's JSON on stdin for deny patterns and exits 2 to block. A hook runs on every tool
call, so a check that starts in about a millisecond instead of booting Node is where the startup
number actually pays.
The common thread
Line the three up by what each one uses to check the agent:
| Tool | Check | Who runs it | What it can prove |
|---|---|---|---|
| Genex self-edit | tsc --noEmit and a booted fork | host, every write | the harness still compiles and starts |
| Genex SkillOpt | 3 blind votes on held-out tasks | a judge model | a model prefers the new text |
| Genex game checks | JS probes and a draw-call ceiling | harness, every build | the game runs and responds |
ffmpeg-skill emit() | ffprobe reads a stream | the script, every write | the output is media |
ffmpeg-skill look.py | a contact sheet | the agent's own eyes | the picture is what was asked |
rdsh guard | a deny-pattern scan | a hook, every tool call | a command matches no pattern |
The pattern is the one Code2Skill found at scale: a skill is worth what its
check is worth. The checks a machine can run (a compiler, a probe, a pattern) are cheap and
certain and narrow. The checks that say whether the work is good (a judge, a look at the frame) are
expensive and soft, and all three projects are candid about which is which. Genex writes "a
learning count does not certify a better game" into its own docs; ffmpeg-skill makes the agent
admit when it never looked; rdsh's README states its machine and n. The marketing compresses
those distinctions. The code keeps them.