~/satyajit

Genex, ffmpeg-skill and rdsh: three harnesses, read for where they check the agent's work

mdjsonmcp

2026-10-06 · 16 min · explainer · agents · harness · open-source · video · rust

Three repositories went round this week with the same shape. Each takes a coding agent that already exists (Claude Code, Codex, DeepSeek's dsh) and wraps it around one domain: games, video, the agent's own launcher. The posts were short: Genex promised a "self-improving harness for game dev", ffmpeg-skill promised 42 tools that probe first and verify after, and rdsh promised 81x faster startup in 799 KB.

I cloned all three with git clone --depth 1 and read them. I did not install or run any of them. Every number below is labelled: measured (I counted it in the source), reported (the project or its post says so; I did not re-run it) or reasoned (my arithmetic on the other two).

The thread that ties them together is the one I care about. A harness, in the sense of Lilian Weng's framing, is the loop and scaffolding around a model. The interesting part of each of these three is where that scaffolding makes the agent check its own work, and what kind of check it is: a compiler, a blind judge, an ffprobe, a picture, or a benchmark.

genex-games/genex-desktop@f23e1c3 · snapshot 2026-10-06
tracked files
2,222
license
MIT
branch
dev
tests
672 files
source
29.1 MB
commit date
2026-10-05
source by language
TypeScript18.6 MB(1635)HTML9.4 MB(29)JavaScript836.0 kB(133)CSS198.3 kB(9)Shell2.6 kB(2)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at f23e1c3 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

kajisho5/ffmpeg-skill@008333a · snapshot 2026-10-06
tracked files
317
license
MIT
branch
main
tests
29 files
source
2.3 MB
commit date
2026-10-05
source by language
Python2.3 MB(82)Shell13.2 kB(3)JavaScript7.5 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at 008333a — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

sahenjp/rustdsh@8b3a8d5 · snapshot 2026-10-06
tracked files
98
license
MIT
branch
main
tests
5 files
source
478.0 kB
commit date
2026-10-06
source by language
Rust246.3 kB(16)JavaScript143.1 kB(15)Shell45.7 kB(9)HTML36.2 kB(4)PowerShell6.8 kB(3)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at 8b3a8d5 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

Genex: a game studio that edits its own instructions

genex-games/genex-desktop is an Electron app, MIT, copyright genex.games, read at commit f23e1c3. Its README links the same genex.games site the post does, so it is theirs. A person describes a game in a chat on the left; a stage on the right shows the running game, its builds and its assets.

Genex banner: the desktop app with a chat about a motocross game on the left and a grid of generated assets, including Blender renders of a motorbike, on the right
Genex's own banner: chat on the left, the game's generated assets on the right. Note the 'Worked in Unreal' rows; the repository's product overview says Unity is retired and the README lists Unity and Unreal plugins as 'soon' (Genex README banner).

How it plugs into a subscription

There is no model API key in the main path. src/substrate/engines/claude-code.ts drives Claude Code through @anthropic-ai/claude-agent-sdk (^0.3.257 in package.json, measured). Its header comment states the boundary in so many words: the studio does not route a consumer subscription through a model API itself; "the SDK spawns Claude Code, which performs its own login and keeps its own credentials in its own config directory." codex.ts does the same for ChatGPT by spawning codex exec --json and reading its JSONL event stream. Both strip ambient keys: an ANTHROPIC_API_KEY in your shell does not count as configured, and CODEX_API_KEY and OPENAI_API_KEY are removed from every child environment, so a stray key cannot flip the bill to per-token (measured from the code and its comments).

Local models go through two more engines: a managed Bonsai 2 27B on Apple Silicon Macs with at least 16 GiB, and Ollama by exact model name (reported, docs/local-models.md). Meshy, Tripo and ElevenLabs are not local. They come through a "Genex Tools" plugin whose manifest reads "made in chat with Genex credits. Coding stays on your own provider." That is the business model, and it is stated plainly.

How it drives Three.js and Blender

The games are browser games on Three.js (three ^0.185.1). The harness does not trust a game to describe itself. Every page the studio serves gets a shim before the game's first line, so window.__studio owns the clock, makes seed(n) reproducible and counts draw calls at the graphics API (reported, docs/harness-runtime.md). Phaser, plain canvas 2D and engine exports are out of scope for that contract.

Blender is a plugin, and its manifest is mostly a command line. The model tool takes a bpy script the agent writes and runs it as blender -b --factory-startup -noaudio --python-exit-code 1 --python wrapper.py -- <script> <model.glb> <render.png> <name>, with a 120,000 ms timeout and a 104,857,600-byte (100 MiB) cap on assets (measured, src/plugins/blender/plugin.json). The wrapper exports model.glb and two renders, render.png and render-front.png, and the backend returns those renders to the agent as images. The skill text attached to the tool says what to do with them: inspect the renders, load the exact GLB path with GLTFLoader, then "verify it in the actual game preview." If Blender is missing, the plugin offers a pinned Blender 5.2.1 download of 346,264,899 bytes, checked against a SHA-256 in the manifest.

What "self-improving" is in the code

This is the claim worth reading for. The harness that builds games is not compiled into the app. It ships as a seed (src/harness-seed/): TypeScript that Node runs by stripping its types, copied into a workspace the in-app agent owns. The docs name four places learning can land, "and they are not interchangeable" (reported, docs/harness-runtime.md):

WhereWhat it holdsIn the seed (measured)
library/checks.jsonexecutable checks on a running game5 checks
library/recipes/*.jsoncraft opinions, each with an optional check44 recipes
skills/*.mdhow an agent works2 skills
prompts/*.mdwhat a role is3 files

The checks are code, not prose. drawcalls-ceiling is the expression __render.drawCalls <= 1000 && __render.triangles <= 400000; player-moved drives the game's own movement keys and asks whether player.x or player.z changed. A recipe pairs a paragraph of intent with a check. characters.feet-on-the-ground says to place a figure by its base rather than its centre, and checks that every object tagged enemy or npc has its bounding-box bottom within 5 cm below or 12 cm above the ground mesh. Its stats in the seed are applied: 0, wins: 0, status candidate (measured). A recipe is retrieved when a check fails or a judge names the defect, "never imposed".

There are two ways those files change.

The agent edits itself. tools/self-tools.ts gives the in-app agent write_own_file, write_skill and install_tool. None of them writes directly. Each goes through guardian.write_self: the host makes a validation fork of the harness, writes the change there, runs a vendored TypeScript 7 tsc --noEmit inside the sandbox with a time limit, boots the fork and asks its healthcheck. Only a pass is written, between two snapshots, with an Activity entry the person can undo. A change may not add type errors; errors the code already had do not block it. The judge rubrics are out of reach: judge/ is write-denied to every agent process, and write_own_file refuses it by name with "judge/ is frozen: the blind critic's rubric is not yours to edit" (measured).

A loop edits the skills between runs. loop/skillopt.ts is a port of Microsoft's SkillOpt, and its header lists what it kept: bounded edits, a ranking, a gate, and a buffer of refused ideas. It mines up to 12 replayable tasks from the run log (facet iterations, refused gates, finished workers, landings), each marked a success or a failure, splits them every other one into a train half and a held-out half, and asks an analyst model for at most 4 edits from four string ops (append, insert_after, replace, delete), never inside a <!-- SLOW_UPDATE --> region. The candidate skill then faces a blind pairwise gate on the held-out tasks: 3 votes, the candidate swapping sides between votes so a judge with a position bias cannot hand it a sweep, accepted only if it wins more than half (measured). A refusal goes into a step buffer of up to 50 so the analyst is shown its dead ideas next time.

Two honest limits sit in the source itself. The split exists because of a failure: "When the analyst saw every task, its candidate restated the very failures the gate then asked about, and all 55 candidates a real install staged won 3/3" (reported, a code comment). And the gate judges text, not games: the docs say skill gates "compare instruction texts against saved task descriptions; they do not execute candidate builds or establish better future game outcomes." So a SkillOpt pass improves what a blind model thinks of the instructions. Whether the next game is better is a separate measurement, and the product copy says as much: "A learning count does not certify a better game." Learning is on by default; automatic application of suggestions is off on fresh installs (reported).

A frame from the Genex launch video: a diff card for skills/director.md replacing 'Place each object by its centre' with 'Place figures by their feet', marked Improved, with checks.json and a feet-on-the-ground recipe card behind it
The launch video's picture of self-improvement: a one-line diff to skills/director.md, marked Improved. The seed's real director.md has no 'Placing things' section; its nine headings run from 'Your seat' to 'Finishing'. The rule shown is the characters.feet-on-the-ground recipe, so read this as an illustration of the mechanism, not a captured run (Genex launch video, about 0:44).

That is a more careful design than the phrase suggested. The agent is allowed to change its own body, but every change passes a compiler and a boot, the yardstick is frozen, and the outer loop keeps an edit only against tasks it did not learn from. It is the same move as the harness-as-generalizer result, where the scaffold rather than the weights is what learns, and it carries the same risk Skill2Env ran into: a skill judged good by a model is not yet a skill shown to help.

ffmpeg-skill: the agent's hands, and a rule to look

kajisho5/ffmpeg-skill is version 2.5.1, MIT, read at 008333a. It is an Agent Skill: a SKILL.md and a scripts/ directory of Python 3.9 standard-library scripts that wrap ffmpeg and ffprobe. npx ffmpeg-skill copies it to ~/.claude/skills/ffmpeg-skill (--cursor, --codex or --all for the others).

Counting the tools. scripts/ holds 42 Python files whose names do not start with an underscore, and scripts/_contract.py describes exactly the same 42, no more and no fewer (measured; I parsed the contract table rather than importing it). By role: 32 execution, 3 analysis, 3 analysis-and-execution (silence, loudness, sync) and 4 verification (check, look, verify, report). 27 tools are flagged visual, and every one of them lists look among the checks to run afterwards (measured). Over MCP, tools/list shows a core 12 by default to save context; the other 30 are still callable by name.

Before and after test pattern clips: the after half is shorter because the quiet stretches were removed
silence.py on synthetic footage: input on the left, output on the right. The repository generates all of its demos from test patterns with demos/build.py, so this shows the mechanics, not a real talk (ffmpeg-skill README, docs/demos/silence_removal.gif).

Probe, plan, run, verify, in the code

The post's "probe first, then process, then verify automatically" is real, but it is two layers, and the split matters.

The first layer is code. Every writing script ends in emit(), which calls verify_output() on the file it just wrote: the file must exist, must not be 0 bytes, and ffprobe must read at least one video or audio stream from it. Otherwise the script dies with kind: output. The JSON result then carries "verified": true only when the file was written, probed and every self-check the tool added passed; a dry run verified nothing. A few tools add their own measurement, such as loudness after a write.

The second layer is instructions. SKILL.md makes the agent probe before planning, run with --dry-run --json to see the exact command lines, compare the output's probed duration, resolution, fps and audio against the request, and, whenever the picture changed, run look.py OUTPUT --tiles 3x2 and view the PNG. "Not finished until Look: names that PNG — a probe cannot see a caption on a face." With no vision, the agent must say so: Look: PATH (pixels not inspected; agent has no image view).

So the code proves the output is media; whether it is the right media is a rule the agent follows. Here is one job through both layers, with the command line each script builds:

ffmpeg-skill · remove the pauses, burn subs.srtclip illustrative · commands from source
Probe what you plan from

Workflow step 1. Every later number comes from this, not from the file name.

agent runs
python3 $S/probe.py talk.mp4 --compact
the script builds
ffprobe -v error -print_format json -show_format -show_streams -show_chapters talk.mp4
what comes back
talk.mp4: 30.000 s, 1920x1080, 30 fps, h264 + aac stereo, VFR suspected: no
0 s30 sred: detected silence · green: kept (5 ranges, 23.800 s)
step 1 of 6

The widget's clip is made up: a 30-second talk with five quiet spans. The commands are not. silence.py asks silencedetect for spans quieter than -35dB for at least 0.6 s, then walks them with keep_ranges(), which keeps a --margin of 0.15 s of silence either side of speech and drops kept pieces shorter than --min-keep 0.2 s. The kept ranges become one expression, between(t,a,b)+between(t,c,d)+…, fed to select for video and aselect for audio, re-encoded in one pass at CRF 18. Moving the margin slider shows the trade: more margin keeps breaths and lengthens the result; less margin tightens it and starts clipping word edges.

Two things the stepper surfaces that the post does not. First, order matters and the skill knows it: silence comes before captions in its chain, and caption.py has an --offset flag but no way to remap cues through silence.py's cut list, so subs.srt has to be timed against the tightened file (or transcribed from it with --transcribe and a local whisper). Second, "verify" for the silence step is structural. The script logs expected ~N s; comparing that against the probe is the agent's job, not the script's.

A test pattern clip on the left and the contact sheet look.py produced from it on the right: twelve tiles with timecodes
look.py's contact sheet, the picture the skill makes the agent open before it reports Done. Synthetic footage, as in all the repository's demos (ffmpeg-skill README, docs/demos/contact_sheet.gif).

The project's own evidence is reported, not re-run: 0 missed gaps on 20 silence.py cases with known gaps, 92 of 92 verification steps on a 10-file real-device corpus, and 72 of 72 agent runs of 24 prompts graded by an independent model, including the visual check run 24 of 24 times the picture changed (README, "Tested on real footage"). The evals/results/ directory keeps those runs as JSON, in iterations numbered up to 25, which is the right habit even if I did not check the grades.

Compared with the launch-video skills, which hand ffmpeg a finished frame sequence at the end, this is the inverse design: the model never writes a filter graph at all. "Nothing runs through a shell; no filter graph is accepted from the caller." The skill also tells the agent not to fall back to raw ffmpeg when no script covers a request, because that "bypasses every guarantee this skill makes."

rdsh: a fast path that mostly doesn't run the agent

sahenjp/rustdsh builds a binary called rdsh, version 0.1.3, MIT, read at 8b3a8d5: 7,214 lines of Rust under src/ (measured). It launches dsh, DeepSeek Harness, the plugin-everything agent harness that deepseek-ai/deepseek-harness ships on npm as @deepseek-ai/dsh and that needs Node ^22.19.0 || >=24.0.0 (measured, its package.json).

What rdsh replaces is the command line in front of dsh, not dsh. Its architecture note is one diagram: argv goes either to a native fast path (tokens, search, compact, sessions, guard, auth, doctor, a status page) or, for anything agentic, to "passthrough: verbatim exec of dsh-orig". On Unix that is a real exec(3): rdsh tui replaces itself with the original dsh, and Node boots exactly as it would have. Installed with --as-dsh, the binary shadows the dsh name and keeps the original as dsh-orig.

The rdsh dashboard page: six cards for doctor, token estimate, prune, sessions, skills and a bench button comparing rdsh and dsh startup
rdsh serve's status page, rendered from the repository's src/ui.html with scripts disabled, so every box shows its empty state. Each card calls a native subcommand; none of them starts an agent (rustdsh, src/ui.html).

What the benchmarks measure

The post's numbers (reported): startup 81x, memory 1/23, search 3.4x, 799 KB against a Node tree of about 508 MB. The README has since moved: about 0.90 ms against 88 ms for startup, which it calls about 98x; 2.9 MB against 66 MB peak memory; search 17 ms against 41 ms, about 2.4x; one binary of about 806 KB (reported, Linux x86_64, median of n=5).

The method is in src/main.rs, and it answers a narrower question than the post implies:

"Slim mode" also does less than its name. slim_env() sets six RDSH_* variables, and the file's own comment says upstream dsh reads none of them; I found no RDSH_ string anywhere in the upstream checkout (measured). The one variable that reaches Node is NODE_COMPILE_CACHE, which caches compiled V8 code on Node 22.1 and later. That can shorten a real dsh boot. The README does not measure by how much.

The repository's docs have drifted from its code in small ways: the README says 26 unit tests and "only three dependencies", while src/ holds 62 #[test] functions and Cargo.toml lists four crates, getrandom included (measured). None of that is a defect in the binary. It is what a project that moved from 81x to 98x in a few days looks like.

The useful part for an agent is small and honest. rdsh guard is a PreToolUse hook command: it scans the hook's JSON on stdin for deny patterns and exits 2 to block. A hook runs on every tool call, so a check that starts in about a millisecond instead of booting Node is where the startup number actually pays.

The common thread

Line the three up by what each one uses to check the agent:

ToolCheckWho runs itWhat it can prove
Genex self-edittsc --noEmit and a booted forkhost, every writethe harness still compiles and starts
Genex SkillOpt3 blind votes on held-out tasksa judge modela model prefers the new text
Genex game checksJS probes and a draw-call ceilingharness, every buildthe game runs and responds
ffmpeg-skill emit()ffprobe reads a streamthe script, every writethe output is media
ffmpeg-skill look.pya contact sheetthe agent's own eyesthe picture is what was asked
rdsh guarda deny-pattern scana hook, every tool calla command matches no pattern

The pattern is the one Code2Skill found at scale: a skill is worth what its check is worth. The checks a machine can run (a compiler, a probe, a pattern) are cheap and certain and narrow. The checks that say whether the work is good (a judge, a look at the frame) are expensive and soft, and all three projects are candid about which is which. Genex writes "a learning count does not certify a better game" into its own docs; ffmpeg-skill makes the agent admit when it never looked; rdsh's README states its machine and n. The marketing compresses those distinctions. The code keeps them.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Genex, ffmpeg-skill and rdsh: three harnesses, read for where they check the agent's work", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026agenttoolsweek,
  author = {Satyajit Ghana},
  title  = {Genex, ffmpeg-skill and rdsh: three harnesses, read for where they check the agent's work},
  url    = {https://ai.thesatyajit.com/articles/agent-tools-week},
  year   = {2026}
}
share