# Genex, ffmpeg-skill and rdsh: three harnesses, read for where they check the agent's work

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/agent-tools-week
> date: 2026-10-06
> tags: explainer, agents, harness, open-source, video, rust

Three repositories went round this week with the same shape. Each takes a coding agent that already
exists (Claude Code, Codex, DeepSeek's `dsh`) and wraps it around one domain: games, video, the
agent's own launcher. The posts were short:
[Genex](https://x.com/genex_games/status/2107148754633507237) promised a "self-improving harness for
game dev", [ffmpeg-skill](https://x.com/bkdgiffug/status/2106505321774526581) promised 42 tools that
probe first and verify after, and [rdsh](https://x.com/sahenjpn/status/2106636746658308369) promised
81x faster startup in 799 KB.

I cloned all three with `git clone --depth 1` and read them. I did not install or run any of them.
Every number below is labelled: **measured** (I counted it in the source), **reported** (the project
or its post says so; I did not re-run it) or **reasoned** (my arithmetic on the other two).

The thread that ties them together is the one I care about. A harness, in the sense of
[Lilian Weng's framing](/articles/agent-harness), is the loop and scaffolding around a model. The
interesting part of each of these three is where that scaffolding makes the agent check its own
work, and what kind of check it is: a compiler, a blind judge, an `ffprobe`, a picture, or a
benchmark.

<RepoCard repo="genex-games/genex-desktop" />

<RepoCard repo="kajisho5/ffmpeg-skill" />

<RepoCard repo="sahenjp/rustdsh" />

## Genex: a game studio that edits its own instructions

[`genex-games/genex-desktop`](https://github.com/genex-games/genex-desktop) is an Electron app, MIT,
copyright `genex.games`, read at commit `f23e1c3`. Its README links the same `genex.games` site the
post does, so it is theirs. A person describes a game in a chat on the left; a stage on the right
shows the running game, its builds and its assets.

<Figure
  src="https://ai.thesatyajit.com/articles/agent-tools-week/fig1.jpg"
  alt="Genex banner: the desktop app with a chat about a motocross game on the left and a grid of generated assets, including Blender renders of a motorbike, on the right"
  caption="Genex's own banner: chat on the left, the game's generated assets on the right. Note the 'Worked in Unreal' rows; the repository's product overview says Unity is retired and the README lists Unity and Unreal plugins as 'soon' (Genex README banner)."
/>

### How it plugs into a subscription

There is no model API key in the main path. `src/substrate/engines/claude-code.ts` drives Claude
Code through `@anthropic-ai/claude-agent-sdk` (`^0.3.257` in `package.json`, measured). Its header
comment states the boundary in so many words: the studio does not route a consumer subscription
through a model API itself; "the SDK spawns Claude Code, which performs its own login and keeps its
own credentials in its own config directory." `codex.ts` does the same for ChatGPT by spawning
`codex exec --json` and reading its JSONL event stream. Both strip ambient keys: an
`ANTHROPIC_API_KEY` in your shell does not count as configured, and `CODEX_API_KEY` and
`OPENAI_API_KEY` are removed from every child environment, so a stray key cannot flip the bill to
per-token (measured from the code and its comments).

Local models go through two more engines: a managed Bonsai 2 27B on Apple Silicon Macs with at
least 16 GiB, and Ollama by exact model name (reported, `docs/local-models.md`). Meshy, Tripo and
ElevenLabs are not local. They come through a "Genex Tools" plugin whose manifest reads "made in
chat with Genex credits. Coding stays on your own provider." That is the business model, and it is
stated plainly.

### How it drives Three.js and Blender

The games are browser games on Three.js (`three` `^0.185.1`). The harness does not trust a game to
describe itself. Every page the studio serves gets a shim before the game's first line, so
`window.__studio` owns the clock, makes `seed(n)` reproducible and counts draw calls at the graphics
API (reported, `docs/harness-runtime.md`). Phaser, plain canvas 2D and engine exports are out of
scope for that contract.

Blender is a plugin, and its manifest is mostly a command line. The `model` tool takes a `bpy`
script the agent writes and runs it as `blender -b --factory-startup -noaudio --python-exit-code 1
--python wrapper.py -- <script> <model.glb> <render.png> <name>`, with a 120,000 ms timeout and a
104,857,600-byte (100 MiB) cap on assets (measured, `src/plugins/blender/plugin.json`). The
wrapper exports `model.glb` and two renders, `render.png` and `render-front.png`, and the backend
returns those renders to the agent as images. The skill text attached to the tool says what to do
with them: inspect the renders, load the exact GLB path with `GLTFLoader`, then "verify it in the
actual game preview." If Blender is missing, the plugin offers a pinned Blender 5.2.1 download of
346,264,899 bytes, checked against a SHA-256 in the manifest.

### What "self-improving" is in the code

This is the claim worth reading for. The harness that builds games is not compiled into the app. It
ships as a seed (`src/harness-seed/`): TypeScript that Node runs by stripping its types, copied into
a workspace the in-app agent owns. The docs name four places learning can land, "and they are not
interchangeable" (reported, `docs/harness-runtime.md`):

| Where | What it holds | In the seed (measured) |
|---|---|---|
| `library/checks.json` | executable checks on a running game | 5 checks |
| `library/recipes/*.json` | craft opinions, each with an optional check | 44 recipes |
| `skills/*.md` | how an agent works | 2 skills |
| `prompts/*.md` | what a role is | 3 files |

The checks are code, not prose. `drawcalls-ceiling` is the expression
`__render.drawCalls <= 1000 && __render.triangles <= 400000`; `player-moved` drives the game's own
movement keys and asks whether `player.x` or `player.z` changed. A recipe pairs a paragraph of
intent with a check. `characters.feet-on-the-ground` says to place a figure by its base rather than
its centre, and checks that every object tagged `enemy` or `npc` has its bounding-box bottom within
5 cm below or 12 cm above the ground mesh. Its stats in the seed are `applied: 0, wins: 0`, status
`candidate` (measured). A recipe is retrieved when a check fails or a judge names the defect, "never
imposed".

There are two ways those files change.

**The agent edits itself.** `tools/self-tools.ts` gives the in-app agent `write_own_file`,
`write_skill` and `install_tool`. None of them writes directly. Each goes through
`guardian.write_self`: the host makes a validation fork of the harness, writes the change there,
runs a vendored TypeScript 7 `tsc --noEmit` inside the sandbox with a time limit, boots the fork and
asks its healthcheck. Only a pass is written, between two snapshots, with an Activity entry the
person can undo. A change may not add type errors; errors the code already had do not block it. The
judge rubrics are out of reach: `judge/` is write-denied to every agent process, and
`write_own_file` refuses it by name with "judge/ is frozen: the blind critic's rubric is not yours
to edit" (measured).

**A loop edits the skills between runs.** `loop/skillopt.ts` is a port of Microsoft's SkillOpt, and
its header lists what it kept: bounded edits, a ranking, a gate, and a buffer of refused ideas. It
mines up to 12 replayable tasks from the run log (facet iterations, refused gates, finished workers, landings), each marked a success or a failure,
splits them every other one into a train half and a held-out half, and asks an analyst model for at
most 4 edits from four string ops (`append`, `insert_after`, `replace`, `delete`), never inside a
`<!-- SLOW_UPDATE -->` region. The candidate skill then faces a blind pairwise gate on the held-out
tasks: 3 votes, the candidate swapping sides between votes so a judge with a position bias cannot
hand it a sweep, accepted only if it wins more than half (measured). A refusal goes into a step
buffer of up to 50 so the analyst is shown its dead ideas next time.

Two honest limits sit in the source itself. The split exists because of a failure: "When the
analyst saw every task, its candidate restated the very failures the gate then asked about, and all
55 candidates a real install staged won 3/3" (reported, a code comment). And the gate judges text,
not games: the docs say skill gates "compare instruction texts against saved task descriptions; they
do not execute candidate builds or establish better future game outcomes." So a SkillOpt pass
improves what a blind model thinks of the instructions. Whether the next game is better is a
separate measurement, and the product copy says as much: "A learning count does not certify a
better game." Learning is on by default; automatic application of suggestions is off on fresh
installs (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/agent-tools-week/fig2.jpg"
  alt="A frame from the Genex launch video: a diff card for skills/director.md replacing 'Place each object by its centre' with 'Place figures by their feet', marked Improved, with checks.json and a feet-on-the-ground recipe card behind it"
  caption="The launch video's picture of self-improvement: a one-line diff to skills/director.md, marked Improved. The seed's real director.md has no 'Placing things' section; its nine headings run from 'Your seat' to 'Finishing'. The rule shown is the characters.feet-on-the-ground recipe, so read this as an illustration of the mechanism, not a captured run (Genex launch video, about 0:44)."
/>

That is a more careful design than the phrase suggested. The agent is allowed to change its own
body, but every change passes a compiler and a boot, the yardstick is frozen, and the outer loop
keeps an edit only against tasks it did not learn from. It is the same move as
[the harness-as-generalizer result](/articles/harness-compositional-generalization), where the
scaffold rather than the weights is what learns, and it carries the same risk
[Skill2Env](/articles/skill2env) ran into: a skill judged good by a model is not yet a skill shown
to help.

## ffmpeg-skill: the agent's hands, and a rule to look

[`kajisho5/ffmpeg-skill`](https://github.com/kajisho5/ffmpeg-skill) is version 2.5.1, MIT, read at
`008333a`. It is an [Agent Skill](https://docs.anthropic.com/en/docs/agents-and-tools/agent-skills):
a `SKILL.md` and a `scripts/` directory of Python 3.9 standard-library scripts that wrap `ffmpeg`
and `ffprobe`. `npx ffmpeg-skill` copies it to `~/.claude/skills/ffmpeg-skill` (`--cursor`,
`--codex` or `--all` for the others).

**Counting the tools.** `scripts/` holds 42 Python files whose names do not start with an
underscore, and `scripts/_contract.py` describes exactly the same 42, no more and no fewer
(measured; I parsed the contract table rather than importing it). By role: 32 execution, 3
analysis, 3 analysis-and-execution (`silence`, `loudness`, `sync`) and 4 verification (`check`,
`look`, `verify`, `report`). 27 tools are flagged `visual`, and every one of them lists `look` among
the checks to run afterwards (measured). Over MCP, `tools/list` shows a core 12 by default to save
context; the other 30 are still callable by name.

<Figure
  src="https://ai.thesatyajit.com/articles/agent-tools-week/fig3.gif"
  alt="Before and after test pattern clips: the after half is shorter because the quiet stretches were removed"
  caption="silence.py on synthetic footage: input on the left, output on the right. The repository generates all of its demos from test patterns with demos/build.py, so this shows the mechanics, not a real talk (ffmpeg-skill README, docs/demos/silence_removal.gif)."
/>

### Probe, plan, run, verify, in the code

The post's "probe first, then process, then verify automatically" is real, but it is two layers,
and the split matters.

The first layer is code. Every writing script ends in `emit()`, which calls `verify_output()` on
the file it just wrote: the file must exist, must not be 0 bytes, and `ffprobe` must read at least
one video or audio stream from it. Otherwise the script dies with `kind: output`. The JSON result
then carries `"verified": true` only when the file was written, probed and every self-check the
tool added passed; a dry run verified nothing. A few tools add their own measurement, such as
loudness after a write.

The second layer is instructions. `SKILL.md` makes the agent probe before planning, run with
`--dry-run --json` to see the exact command lines, compare the output's probed duration,
resolution, fps and audio against the request, and, whenever the picture changed, run
`look.py OUTPUT --tiles 3x2` and view the PNG. "Not finished until `Look:` names that PNG — a probe
cannot see a caption on a face." With no vision, the agent must say so:
`Look: PATH (pixels not inspected; agent has no image view)`.

So the code proves the output is media; whether it is the right media is a rule the agent follows.
Here is one job through both layers, with the command line each script builds:

<ProbeActVerify />

The widget's clip is made up: a 30-second talk with five quiet spans. The commands are not.
`silence.py` asks `silencedetect` for spans quieter than `-35dB` for at least `0.6` s, then walks
them with `keep_ranges()`, which keeps a `--margin` of 0.15 s of silence either side of speech and
drops kept pieces shorter than `--min-keep` 0.2 s. The kept ranges become one expression,
`between(t,a,b)+between(t,c,d)+…`, fed to `select` for video and `aselect` for audio, re-encoded in
one pass at CRF 18. Moving the margin slider shows the trade: more margin keeps breaths and
lengthens the result; less margin tightens it and starts clipping word edges.

Two things the stepper surfaces that the post does not. First, order matters and the skill knows
it: silence comes before captions in its chain, and `caption.py` has an `--offset` flag but no way to
remap cues through `silence.py`'s cut list, so `subs.srt` has to be timed against the tightened
file (or transcribed from it with `--transcribe` and a local whisper). Second, "verify" for the
silence step is structural. The script logs `expected ~N s`; comparing that against the probe is the
agent's job, not the script's.

<Figure
  src="https://ai.thesatyajit.com/articles/agent-tools-week/fig4.gif"
  alt="A test pattern clip on the left and the contact sheet look.py produced from it on the right: twelve tiles with timecodes"
  caption="look.py's contact sheet, the picture the skill makes the agent open before it reports Done. Synthetic footage, as in all the repository's demos (ffmpeg-skill README, docs/demos/contact_sheet.gif)."
/>

The project's own evidence is reported, not re-run: 0 missed gaps on 20 `silence.py` cases with
known gaps, 92 of 92 verification steps on a 10-file real-device corpus, and 72 of 72 agent runs of
24 prompts graded by an independent model, including the visual check run 24 of 24 times the picture
changed (README, "Tested on real footage"). The `evals/results/` directory keeps those runs as JSON,
in iterations numbered up to 25, which is the right habit even if I did not check the grades.

Compared with [the launch-video skills](/articles/one-shot-launch-videos), which hand `ffmpeg` a
finished frame sequence at the end, this is the inverse design: the model never writes a filter
graph at all. "Nothing runs through a shell; no filter graph is accepted from the caller." The skill
also tells the agent not to fall back to raw `ffmpeg` when no script covers a request, because that
"bypasses every guarantee this skill makes."

## rdsh: a fast path that mostly doesn't run the agent

[`sahenjp/rustdsh`](https://github.com/sahenjp/rustdsh) builds a binary called `rdsh`, version
0.1.3, MIT, read at `8b3a8d5`: 7,214 lines of Rust under `src/` (measured). It launches `dsh`,
[DeepSeek Harness](/articles/deepseek-harness), the plugin-everything agent harness that
[`deepseek-ai/deepseek-harness`](https://github.com/deepseek-ai/deepseek-harness) ships on npm as
`@deepseek-ai/dsh` and that needs Node `^22.19.0 || >=24.0.0` (measured, its `package.json`).

What rdsh replaces is the command line in front of `dsh`, not `dsh`. Its architecture note is one
diagram: argv goes either to a native fast path (`tokens`, `search`, `compact`, `sessions`,
`guard`, `auth`, `doctor`, a status page) or, for anything agentic, to "passthrough: verbatim exec
of dsh-orig". On Unix that is a real `exec(3)`: `rdsh tui` replaces itself with the original `dsh`,
and Node boots exactly as it would have. Installed with `--as-dsh`, the binary shadows the `dsh`
name and keeps the original as `dsh-orig`.

<Figure
  src="https://ai.thesatyajit.com/articles/agent-tools-week/fig5.png"
  alt="The rdsh dashboard page: six cards for doctor, token estimate, prune, sessions, skills and a bench button comparing rdsh and dsh startup"
  caption="rdsh serve's status page, rendered from the repository's src/ui.html with scripts disabled, so every box shows its empty state. Each card calls a native subcommand; none of them starts an agent (rustdsh, src/ui.html)."
/>

### What the benchmarks measure

The post's numbers (reported): startup 81x, memory 1/23, search 3.4x, 799 KB against a Node tree of
about 508 MB. The README has since moved: about 0.90 ms against 88 ms for startup, which it calls
about 98x; 2.9 MB against 66 MB peak memory; search 17 ms against 41 ms, about 2.4x; one binary of
about 806 KB (reported, Linux x86_64, median of n=5).

The method is in `src/main.rs`, and it answers a narrower question than the post implies:

- **Startup** is `rdsh --version` against `dsh --version`, each spawned 5 times and the median
  taken. rdsh's `--version` is printed by `clap` and exits; the comment in the source says so
  ("--version/-V is served by clap itself"). `dsh --version` boots a Node runtime. So the 98x is a
  real measurement of how long it takes to print a version string. It is not how long an agent
  session takes to start, because `rdsh tui` still execs `dsh` and pays Node's boot plus about a
  millisecond (reasoned).
- **Search, tokens and sessions** compare rdsh with rdsh: "before" is a build of rdsh's previous
  HEAD, "after" is the working tree, outputs diffed for equality. The 2.4x is an optimisation inside
  rdsh. There is no search-against-`dsh` row.
- **Memory** is the same `--version` comparison under `/usr/bin/time -v`.
- **Size** compares rdsh's binary with `dsh`'s Node install, but rdsh needs that Node install to do
  anything agentic, so the two do not substitute (reasoned).

"Slim mode" also does less than its name. `slim_env()` sets six `RDSH_*` variables, and the file's
own comment says upstream `dsh` reads none of them; I found no `RDSH_` string anywhere in the
upstream checkout (measured). The one variable that reaches Node is `NODE_COMPILE_CACHE`, which
caches compiled V8 code on Node 22.1 and later. That can shorten a real `dsh` boot. The README does
not measure by how much.

The repository's docs have drifted from its code in small ways: the README says 26 unit tests and
"only three dependencies", while `src/` holds 62 `#[test]` functions and `Cargo.toml` lists four
crates, `getrandom` included (measured). None of that is a defect in the binary. It is what a
project that moved from 81x to 98x in a few days looks like.

The useful part for an agent is small and honest. `rdsh guard` is a `PreToolUse` hook command:
it scans the hook's JSON on stdin for deny patterns and exits 2 to block. A hook runs on every tool
call, so a check that starts in about a millisecond instead of booting Node is where the startup
number actually pays.

## The common thread

Line the three up by what each one uses to check the agent:

| Tool | Check | Who runs it | What it can prove |
|---|---|---|---|
| Genex self-edit | `tsc --noEmit` and a booted fork | host, every write | the harness still compiles and starts |
| Genex SkillOpt | 3 blind votes on held-out tasks | a judge model | a model prefers the new text |
| Genex game checks | JS probes and a draw-call ceiling | harness, every build | the game runs and responds |
| ffmpeg-skill `emit()` | `ffprobe` reads a stream | the script, every write | the output is media |
| ffmpeg-skill `look.py` | a contact sheet | the agent's own eyes | the picture is what was asked |
| rdsh `guard` | a deny-pattern scan | a hook, every tool call | a command matches no pattern |

The pattern is the one [Code2Skill](/articles/code2skill) found at scale: a skill is worth what its
check is worth. The checks a machine can run (a compiler, a probe, a pattern) are cheap and
certain and narrow. The checks that say whether the work is good (a judge, a look at the frame) are
expensive and soft, and all three projects are candid about which is which. Genex writes "a
learning count does not certify a better game" into its own docs; ffmpeg-skill makes the agent
admit when it never looked; rdsh's README states its machine and `n`. The marketing compresses
those distinctions. The code keeps them.
