REA, Reverse Engineer Anything: 139 MCP tools between your agent and a decompiler, each with a receipt
mdjsonmcp2026-10-08 · 30 min · mcp · agents · security · developer-tools · agentic-coding · reverse-engineering
Why read this
Notabletop 60%REA's 139 tools counted from the package, both decompiler bridges read, the DX-Ball case walked, and a measured blind spot on minified Electron code.
- Original analysis
- Runs on a laptop CPU
- A new technique
Developer tools & infraMITPractitioner tool
How this was scored
- Is it new?
- 2 of 3: A real new idea, method or capability
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 62 of 100, ranked 221 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
The pitch on REA's README is one line: "See a feature you like. Understand how it works, down to the binary level." You point your coding agent at an app you have installed, ask how a feature works, and the agent drives a disassembler for you, explains the code it finds, and writes a version for your project.
I expected a thin wrapper. There are already MCP servers that put Ghidra or IDA in front of a model, and most of them are a few hundred lines that forward decompile_function to the engine and paste the result into the chat. REA is not that. The published package registers 139 tools. The source tree has about 167,000 lines of TypeScript outside its tests and another 119,000 lines of tests, a Python script that runs inside Hopper, a 3,321-line Java script that runs inside Ghidra, and a Node-API addon for Windows. And almost every tool returns the same thing: a result plus a content-addressed record of where it came from.
That record is the interesting part. A decompiler is a guesser, and an LLM is a guesser on top of it. REA's bet is that if every observation carries its target digest, its addresses, its provider and its stated limitations, the agent and the person reading the agent's answer can tell what the binary says from what the model thinks. I wanted to know whether that bet holds, so I read the code, walked through their best case study, and ran the part that runs without a commercial disassembler on an app I wrote.
- license
- MIT
- branch
- main
- tests
- 806 files
- source
- 11.7 MB
- commit date
- 2026-10-08
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-08 at bc4aa46 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
What it is, in one picture
REA is a Node program. Run as rea mcp (or through npx rea-agents), it is a stdio MCP server; run as rea <command>, it is a CLI with the same workflows behind it. npx rea-agents setup finds the agents installed on your machine (Claude Code, Codex, Cursor, Gemini CLI and eight more by the skill's list), prints the config changes it wants to make, backs up what it overwrites, and on approval registers the server and drops a workflow skill into each agent. It can also install Hopper, again only after you say yes.

The "analysis tools" box in that figure hides most of the engineering. Underneath it there are three kinds of thing:
- Deep native providers: Hopper, Ghidra or IDA, each driven through a bridge that runs inside the engine. You bring Ghidra and IDA; setup can fetch Hopper.
- REA's own readers, written in TypeScript: ASAR archives, JavaScript and HTML parsed with Babel, PE/CLI metadata and CIL for .NET, ZIP/IPA/APK/MSIX/DMG inventories, plists, Mach-O headers.
- Wrapped third-party tools you supply: JADX for Android, Binwalk and Unblob for firmware, pwntools and pwndbg for ELF layout and core files, Wakaru for unbundling JavaScript, EVMole for contract bytecode, Chrome over the DevTools protocol and Playwright for websites.
Nothing leaves the machine except what your agent sends to its own model provider; the README's FAQ says exactly that, and it is the honest version of "runs locally".
139 tools, counted
The README links a "tool catalog", but the catalog is build-generated and not in the repository. So I installed rea-agents@6.0.0 from npm into a scratch directory with --ignore-scripts, imported TOOL_CONTRACTS from the package's dist/contracts/toolContracts.js, and converted each input and output schema to JSON Schema. The order of that array is fixed in src/contracts/toolContracts.ts:24-44, and it is built from nineteen contract files, which give a natural grouping by target.
Several primitives joined into one answer: the function dossier, feature traces, call paths, Objective-C and Swift metadata.
analyze_functionNative: composed analysisBuild a dossier for one native function identified by symbol or provider-returned address.
procedureresultevidence_idevidenceresult, evidence_id and evidence; only 5 declare themselves read-only, because recording Evidence counts as a change to the session.Here are the counts, grouped by what you point them at:
| Target | Tools | What they are |
|---|---|---|
| Native binaries, disassembler primitives | 41 | list_procedures, procedure_pseudo_code, xrefs, read_bytes, set_comment… one call each on the open database |
| Native binaries, composed | 14 | analyze_function, trace_feature, trace_call_path, batch_decompile, Objective-C and Swift metadata |
| Native host (macOS) | 8 | Mach-O and signatures, accessibility trees, observe_native_calls under LLDB |
| Offline ELF and crash cores | 2 | pwntools layout, recorded crashes |
| Packages and bundles | 6 | inventory, extraction, Interface Builder, asset catalogs, dylib resolution |
| .NET assemblies | 7 | metadata, members and CIL, cross-build comparison, imported decompiler output |
| Android | 5 | JADX package, class, method and reference queries |
| Firmware | 2 | Binwalk/Unblob regions and extraction |
| EVM bytecode | 1 | selectors and mutability via EVMole |
| JavaScript and Electron | 8 | analyze_javascript_application, source recovery, V8 and Electron observation |
| Websites and captures | 15 | CDP page inspection, Playwright scenarios, script export, source maps, HAR |
| Cross-target workflows | 9 | trace_application_feature, version comparison, reconstruction ledgers |
| Session, Evidence, unknowns | 21 | open_binary, close_binary, Evidence bundles, record_unknown, verify_reconstruction |
Two things stand out. The 41 primitives are a disassembler's vocabulary: they existed first, when this project was a Hopper bridge (its changelog's 0.2.0 entry is "rename package and CLI to REA"), and most of them map straight onto an operation name in the Hopper bridge's dispatcher. The other 98 are where the project went after that, and most of them are not about native code at all.
The other is the output shape. 126 of the 139 tools return { result, evidence_id, evidence }. Twelve session tools return a bare result, and one returns a reconstruction-coverage verdict. Only 5 tools set the MCP readOnlyHint, and docs/mcp-contracts.md explains why: "an analysis call that records additive Evidence is marked non-read-only even when it leaves the target unchanged." Pedantic, and correct. A client that auto-approves read-only tools will ask about almost all of these.
Inside Hopper: a script that serves a socket
Hopper is a commercial disassembler for macOS and Linux with a Python scripting console. It has no server mode. So REA starts Hopper with a script and makes the script the server.
src/hopper/BridgeLauncher.ts writes a small bootstrap.py into a private session directory with mode 0600, holding the socket path, a random token, a run id and the target path:
// src/hopper/BridgeLauncher.ts:326-330
`REA_SOCKET = ${JSON.stringify(session.socketPath)}`,
`REA_TOKEN = ${JSON.stringify(session.token)}`,
`REA_RUN_ID = ${JSON.stringify(session.runId)}`,
`REA_TARGET_PATH = ${JSON.stringify(options.targetPath)}`,
`REA_OWNS_PROCESS_LIFETIME = ${ownsProcessLifetime ? "True" : "False"}`,and then launches Hopper on the target with that script:
// src/hopper/BridgeLauncher.ts:162-170
const action =
this.options.targetKind === "database" ? "--database" : "--executable";
const argumentsForTarget = [
...this.options.loaderArgs,
"--analysis",
"-Y",
bootstrapPath,
action,
this.options.targetPath,
];The bootstrap runs bridge/hopper_bridge.py (1,020 lines) on Hopper's own Python thread. Its docstring carries the one hard-won rule of the whole file: "Keep all Hopper API access on this thread: moving dispatch to a worker can deadlock Hopper." The bridge binds a Unix socket, chmods it to 0600, accepts exactly one client, and reads newline-delimited JSON. Every request must have exactly four keys and the right token, compared in constant time:
# bridge/hopper_bridge.py:912-925
if not isinstance(request, dict) or set(request) != {"id", "token", "method", "params"}:
raise InvalidRequestError("Invalid bridge request shape")
...
if (
not isinstance(request["token"], str) or not request["token"].isascii()
or not hmac.compare_digest(request["token"], REA_TOKEN)
):
raise PermissionError("Invalid bridge capability")SECURITY.md calls this the security boundary, and is plain about its limits: a random capability token plus a current-user socket "is not a sandbox and does not protect against malicious processes already running as the same operating-system user." Opening an untrusted binary means Hopper parses it with your permissions.
The screenshot on REA's README is this mechanism caught in the act. Look at the console at the bottom.

Hopper's background analysis finishes, then "Executing Python script /tmp/rea-sAWZXd/bootstrap.py". From that moment the agent's analyze_function calls arrive over the socket, and the bridge answers them with Hopper's own API.
The bridge's _analyze_function (bridge/hopper_bridge.py:528-621) is worth reading because it is the dossier every native investigation leans on. It walks the procedure's basic blocks and successors, takes Hopper's pseudocode and assembly, collects callers and callees, every reference into and out of the function's instructions, the strings and names those references hit, and the comments. Then it says what it could not do:
# bridge/hopper_bridge.py:614-619
"limitations": [
"Hopper's public Python API does not classify reference kinds.",
"Hopper's public Python API does not expose equivalent external or thunk classification in this dossier.",
"Unresolved indirect calls without reported target addresses are not represented as call edges.",
"Pseudocode and assembly are provider-specific representations, not original source.",
],The last line is the one an agent most needs to hear and least often does.
There is one more Hopper detail I did not expect. On Linux, REA runs Hopper's demo build headless on a private Xvfb display, and something has to choose the demo in Hopper's registration dialog. scripts/hopper-demo-x11.py does it with a synthetic XTest click at a fixed offset from the dialog's corner (DEMO_CLICK_LEFT = 116, DEMO_CLICK_BOTTOM = 22), and only for two Hopper builds whose SHA-256 it pins. It is careful engineering, and it is also automation of a commercial product's demo gate, which I come back to below.
Ghidra: a postScript that never exits
Ghidra is free (Apache-2.0) and has a proper headless mode, so REA uses it the way Ghidra intends to be scripted. ghidraHeadlessArguments builds one analyzeHeadless call:
// src/ghidra/GhidraLauncher.ts:312-358, abridged
options.projectRoot, "rea-project",
"-import", options.targetPath,
"-readOnly",
"-deleteProject",
"-log", options.ghidraLogPath,
"-scriptlog", options.scriptLogPath,
"-scriptPath", dirname(options.bridgeScriptPath),
"-postScript", options.bridgeScriptPath, options.descriptorPath,-readOnly and -deleteProject mean a throwaway project in a private directory; REA never opens your own Ghidra projects. -postScript runs bridge/ghidra/ReaGhidraBridge.java after auto-analysis. The script extends Ghidra's HeadlessScript (ReaGhidraBridge.java:101), reads a session descriptor and deletes it (:160-174), binds a Unix domain socket with rw------- permissions (:232-234), and then does not return. It sits in a switch over operation names (:352-377), with the same names as the Hopper bridge: procedure_pseudo_code, xrefs, read_bytes, analyze_function. That shared vocabulary is what lets the MCP side stay provider-neutral.
At 3,321 lines the Ghidra bridge does more than Hopper's. It decompiles with DecompInterface, walks Ghidra's ClangCaseTokens to recover switch cases (and refuses to guess when two typed labels disagree), and exposes function annotation, which on Ghidra edits only the ephemeral database: the executable's bytes never change. An extension point loads Washi1337's ghidra-nativeaot to recover metadata from .NET NativeAOT binaries, and a -preScript handles DOS .COM files at segment 1000:0100. That last one exists because two of REA's showcases are a 1996 Windows game and a PC-98 DOS game.
IDA is the third provider, and the thinnest: REA adapts the existing ida-pro-mcp server rather than writing its own plugin. native/ holds the Windows addon, about 1,100 lines of C++ that admit input files by NTFS file id, create private directories with protected DACLs and assign child processes to Job Objects. The package refuses to load it unless its version, ABI, architecture and SHA-256 match.
Electron and JavaScript: a graph from inert syntax
For JavaScript, REA does not need any engine. analyze_javascript_application reads an extracted app directory or an .asar (through @electron/asar), parses every script and HTML file with @babel/parser, and never executes any of it. From the AST it builds what the docs call a JavaScript Application Graph: package, main, preload and renderer code, BrowserWindows, contextBridge APIs, IPC channels and handlers, imports, source maps, native add-ons, with typed edges (loads, exposes, invokes, handles, imports, maps_to). Node and graph ids are SHA-256 digests of the normalized content, so the same input gives the same graph id every time.
This is the one part I could run, because it needs only Node. I wrote a five-file Electron app, a note search with a preload bridge: main.js creates a BrowserWindow with preload.js and handles notes:search; preload.js exposes window.notes.search, which invokes that channel; renderer.js calls it. Then:
npx rea-agents@6.0.0 analyze-javascript-application /abs/path/to/app --jsonIt took under 3 seconds and found the whole chain. package.json loads main.js; a BrowserWindow at main.js 6:14 loads preload.js; preload.js exposes the notes API and invokes notes:search at 3:21; the handler at main.js line 12 handles it. Every edge carries a source range, a confidence and the label inferred, and the result's limitations say why: "Electron relationships are derived from inert syntax; runtime registration, reachability, defaults, and enforcement remain unproven." It also listed 45 unknowns, mostly dynamic calls it would not resolve. Running it twice gave the same evidence id, ev_620fc9de….
It did not link renderer.js to the bridge it calls. window.notes.search(...) in the renderer stays an unconnected call; the graph knows the renderer loads the script and knows the bridge exists, but not that one uses the other. And the output for 1,355 bytes of source was 617 KB of compact JSON. On a real app that is a lot of context to hand a model, which is presumably why docs/mcp-contracts.md spends a section on what happens when a result does not fit the client's 10 MiB receive buffer.
Then I minified the app, because no shipped Electron app looks like my readable version.
const { contextBridge, ipcRenderer } = require("electron");
contextBridge.exposeInMainWorld("notes", {
search: (query) => ipcRenderer.invoke("notes:search", query),
});const { app, BrowserWindow, ipcMain } = require("electron");
…
ipcMain.handle("notes:search", (_event, query) => searchNotes(query));The whole chain, each edge with a source range: package.json loads main.js, a BrowserWindow loads preload.js, preload exposes the notes API and invokes notes:search, and main.js line 12 handles it.
With a bundler's output, every Electron count went to zero. The cause is in src/domain/javascript/electronStaticAnalysisBrowser.ts:161-167:
const name = calleeName(node.callee);
const main =
name === "contextBridge.exposeInMainWorld" ||
name.endsWith(".contextBridge.exposeInMainWorld");calleeName (javascriptStaticAnalysisHelpers.ts:570) spells out the callee's member path as text. e.contextBridge.exposeInMainWorld(...) matches. n.exposeInMainWorld(...), where n was destructured from require("electron") as {contextBridge: n}, does not, and esbuild writes exactly that. The IPC matcher works the same way with ipcMain. and ipcRenderer. suffixes. Text matching also cuts the other way: any object a program happens to call ipcMain would match, whether or not it came from Electron. I have not seen that happen, but it follows from the code.
What bothers me is less the miss than its silence. The minified result says it parsed every file with no failures and no truncation. Its limitations are five generic sentences, including "Static paths and relationships may remain unresolved when expressions are dynamic or obfuscated." The readable run had six; the missing one is the caveat about Electron relationships, dropped, presumably, because as far as the analyzer could tell there were none. None of its 47 unknowns says "an Electron binding was renamed here and I stopped following it". An agent reading context_bridge_apis: 0 would reasonably conclude the app has no bridge.
REA's own Notion case study runs into the same wall from the other side, and to its credit says so. On Notion Desktop's minified tab preload, the analyzer found an exposed API with api_status: "dynamic" and an IPC call whose channel was the variable e. Tracing from the real channel name, notion:clipboard:write, found "no graph match", and the page's evidence note says the clipboard route came from reading the wrapper code, "not an automatically paired IPC graph". An honest outcome, and the normal one for production JavaScript, and it means the agent does the last mile with its own eyes.
For bundles that need unpacking first, recover_javascript_sources runs Wakaru (pinned at v1.13.0, bring your own binary) under prlimit, writes the modules to a fresh directory, and records the exact input digest and the executable's SHA-256. For live sites, REA attaches over CDP to a Chrome you started with --remote-debugging-port, and keeps attach-only observation (inspect_web_page, observe_web_session) apart from Playwright scenarios that launch and drive a browser, because the two have different authority. It can even list a page's WebMCP tool registrations, which is the subject of the WindTunnel WebMCP benchmark on this site.
.NET without ILSpy
I assumed the .NET path wrapped ILSpy or dnlib. It does neither. src/dotnet/ is REA's own PE/CLI reader in TypeScript: PE headers, the CLI header, metadata heaps and tables, method bodies and a CIL decoder (ManagedMemberInstructionDecoder.ts, 551 lines). It never loads or executes an assembly. ILSpy appears only as an optional oracle: if you point REA_ILSPY_CMD_PATH at ilspycmd, import_managed_reconstruction will attach its C# to the exact members REA parsed, labelled as inference. The managed-code guide is blunt about the boundary: "it does not unpack .NET single-file hosts, decode IL2CPP metadata, or infer a NativeAOT identity from an ordinary native PE." NativeAOT goes to Ghidra.
The receipt
Every one of those 126 Evidence-returning tools builds the same record. Its schema is in src/domain/evidence.ts:
// src/domain/evidence.ts:96-113
const evidenceBaseSchema = z
.object({
evidence_id: prefixedDigestSchema("ev"),
subject: subjectSchema.nullable(),
provider: providerSchema,
predicate_type: z.string().min(1),
operation: z.string().min(1),
parameters: jsonObjectSchema,
raw_result: jsonValueSchema.nullable(),
normalized_result: jsonValueSchema,
confidence: z.enum(["observed", "derived", "inferred"]),
authority: evidenceAuthoritySchema,
environment: executionEnvironmentSchema.nullable(),
limitations: z.array(z.string()),
locations: z.array(evidenceLocationSchema),
evidence_links: z.array(prefixedDigestSchema("ev")),
})
.strict();The subject carries the target's SHA-256, format and architecture. provider names the engine and its version. locations are addresses, address ranges, file offsets or archive paths. authority says what kind of source it is: the shipped artifact, a controlled replay, a historical reference, an external service, or analyst inference. And the id is not random: computeEvidenceId (evidence.ts:229-230) is ev_ plus a SHA-256 of the canonical JSON of everything except the display path, and parseEvidence rejects a record whose id does not match its content. You cannot quietly edit a result and keep its id.
On top of the records sits an unknowns register: record_unknown, update_unknown with an expected revision, and verify_unknown_resolution, which only accepts "qualifying observed evidence" as a resolution. Comparisons (compare_functions, compare_artifacts, find_changed_behavior) take Evidence as input, not paths, so a claim that two versions differ points at the two records it came from.
The skill REA installs into your agent (skill-src/reverse-engineer-anything/SKILL.md, 216 lines) is how all that reaches the model's behaviour. It is mostly routing ("ASAR or extracted JavaScript/Electron tree: analyze_javascript_application"), and then a few rules that do the real work:
Every conclusion must distinguish observations, inferences, and unknowns. Cite Evidence IDs, preserve limitations and incomplete coverage, and never imply that static analysis observed execution.
and, for anything bigger than one function:
Keep a concise finding ledger linking each conclusion to Evidence IDs, confidence/evidence type, search boundary, and remaining unknowns. [...] Before finishing, revisit the original checklist. Mark each question as answered, partially answered, or unresolved based on its evidence; keep bounded negative searches bounded, and do not describe a broad investigation as complete while required questions remain open.
There is also one line I would underline for anyone pointing an agent at a stranger's software: "Never choose an example app on the user's behalf."
Six MCP prompts (investigate_feature, compare_application_versions, verify_reconstruction, trace_crash, audit_residual_unknowns, prepare_bounded_process_capture) give the same guidance through the protocol for clients that do not load skills.
The limit of all this is structural. The receipt proves what a tool returned. It does not constrain what the agent says about it. The final explanation is prose the model writes, and a skill instruction to "distinguish observations, inferences, and unknowns" is a request, not a check. What the contract does buy is auditability: when an agent says "this function takes the brick's x position", you can ask which Evidence id that came from and go and look. A real improvement over a chat transcript of pasted decompiler output, and still not verification.
The DX-Ball case, step by step
REA's best showcase is a single function from DX-Ball 1.07, the 1996 Windows breakout game: the one that decides how far left or right a brick-break sound is panned. It is small enough to check by hand, and it contains the exact failure that makes people distrust decompilers.

The target is DXBALL.EXE, 158,208 bytes, PE i386, identified by SHA-256. The recorded analysis used REA 4.1.0 with Ghidra 12.1.4 on Linux. Step through it:
Use REA to find how DX-Ball calculates sound panning. Explain the calculation and show the code.
DXBALL.EXE 158,208 bytes PE i386 sha256 756da1ba09edce71…44e70490 DX-Ball 1.07, 1996, English Windows build
The question is about behaviour you can hear. The target is a local copy of the original executable, identified by its digest, which every later Evidence record carries as its subject.
The decompiler step is the one to stop on. Ghidra's pseudocode for 0x406400 is longlong FUN_00406400(void) returning __ftol(): no argument, no arithmetic. An agent handed only that, which is what most decompiler MCP servers hand it, would describe a function that converts nothing to an integer, or worse, invent a purpose for it from its callers. REA's dossier puts the 23 instructions next to the pseudocode, and the instructions show a stack argument at [EBP+8], an FILD, an FMUL by the double at 0x420068, an FSUB of the double at 0x420070, and a second FMUL by the global at 0x4210a0. That disagreement inside one result is what made the agent keep going.

The arithmetic is satisfying once you have all three pieces. The caller computes 20 + 30 × tile_x with an ADD EAX, EAX and two LEAs (×2, ×3, ×5) and an ADD EAX, 0x14. The function multiplies that by 1.5625, which is 1000/640, so a 640-pixel screen spans 0 to 1000; subtracting 500 centres it. At a scale of 1 the result is -500 at the left edge, 0 in the middle and 500 at the right, which is a stereo pan.

Then comes the part REA did not do. The reconstruction project (N0zoM1z0/dx-ball, MIT for its own code) runs the original x86 function against the C for every integer position from 0 to 640 at five scales (0, 0.5, 1, 20 and -1), which is 641 × 5 = 3,205 cases, and rebuilds the C with the pinned VC4.0 compiler until all 63 bytes of the function match. REA supplied the evidence. The oracles that make the result trustworthy live in the project.
I wanted to see why byte-matching needs the original compiler, so I compiled the same C myself:
gcc -m32 -O0 -mfpmath=387 -fno-pic -c pan.c # GCC 13.3.0
objdump -d -M intel pan32.o ; objdump -s -j .rodata -j .data pan32.oWhat the CPU executes. The left column is VC4.0 output from 1996; the right is a modern GCC at -O0.
PUSH EBP / MOV EBP, ESP / SUB ESP, 0xc PUSH EBX / PUSH ESI / PUSH EDI MOV EAX, [EBP+0x8] MOV [EBP-0xc], EAX FILD dword [EBP-0xc] FMUL qword [0x420068] FSUB qword [0x420070] FMUL qword [0x4210a0] CALL 0x41678c ; __ftol POP EDI / POP ESI / POP EBX LEAVE / RET 23 instructions, 63 bytes
push ebp / mov ebp, esp / sub esp, 0x18 fild dword [ebp+0x8] fstp qword [ebp-0x8] fld qword [ebp-0x8] fld qword [.rodata+0] fmulp st(1), st … fsubp, fmulp … fnstcw / or ah, 0xc / fldcw fistp dword [ebp-0x18] ; inline truncation fldcw / mov eax, [ebp-0x18] leave / ret 85 bytes
Same C, different code. VC4.0 calls a runtime helper to truncate a double; GCC switches the x87 rounding mode and stores inline. A byte-exact reconstruction therefore has to be checked against the original compiler, which is what the DX-Ball project's pinned VC4.0 replay does.
My constants came out byte-identical to the ones REA read from DX-Ball: 000000000000f93f and 0000000000407f40 in .rodata, 000000000000f03f in .data. The code did not. GCC's function is 85 bytes against DX-Ball's 63, and where VC4.0 calls the runtime helper __ftol to truncate the double, GCC switches the x87 control word and does an inline fistp. Equivalent behaviour, different bytes. A "63 matching bytes" claim is only possible because the project found and pinned the 1996 compiler.
The DX-Ball repository has moved on since the case study's 7 October checkpoint (55 maintained functions then). Its README at the commit I cloned reports 243 functions present in source, 111,242 differential cases against the original, and 35 byte-exact functions totalling 3,379 bytes, with experimental Windows builds running under Wine. I did not run any of it: it needs the original game files, which the repository does not ship.
One disclosure the case study does not make: N0zoM1z0, who owns both the DX-Ball and the TH04 reconstruction repositories, is REA's second-largest contributor, with 273 commits to the maintainer's 836 according to GitHub's contributor list. The showcases are the team's own work. That does not make them wrong (the evidence is published and checkable), but they are not independent users' reports.
What an agent with a decompiler can and cannot do
DX-Ball is a good case because it is a friendly target: an unobfuscated 1996 binary built by a known compiler, with a function whose purpose you can hear. Most targets are worse in at least one of these ways, and REA's tooling does not change that.
Start with names. The executable had no symbol for 0x406400. dxball_screen_pan, x and dxball_pan_scale are the project's names, and the evidence note says so in a table with two columns: "Saved observation" and "Interpretation used by the reconstruction". That separation is the right habit. An agent that names a stripped function check_license because it references a string containing "license" has written a hypothesis, and REA's record will show the string reference, not the purpose. The risk of a confident wrong purpose is the model's, and REA can only make it easier to catch.
The decompiler output is a guess that looks like source. The DX-Ball pseudocode was wrong in a way that would have sent a pseudocode-only agent off a cliff. REA's answer is to return assembly with pseudocode by default, and the Hopper bridge's last limitation sentence says it outright. Whether the model reads the assembly when the pseudocode looks plausible is up to the model.
Obfuscation and packing defeat static reading, and REA mostly says so only in general terms. The skill's native guide notes that "decompiling an unpacking stub does not recover the unpacked program". The JavaScript analyzer's limitations mention obfuscation. But as my minified Electron run showed, a specific miss is not always reported as a specific unknown. The coverage status for that run was unknown, which is technically honest and practically easy to skim past.
In the end, what survives is what gets tested. The two DX-Ball claims I trust most are the 3,205 differential cases and the 63-byte compiler match, and neither is an LLM output. REA has a tool for this (verify_reconstruction evaluates a typed specification against Evidence and, in its own words, a pass "covers only declared comparable claims, not global equivalence"), but the heavy lifting is running the original code, which needs your own harness.
My summary for someone deciding whether to use it: REA makes an agent much better at gathering the right facts from a binary and much easier to audit. It does not make the agent's interpretation of those facts correct, and it is not designed to.
How mature is it
Young and fast. The repository was created on 14 April 2026; the first npm release, 0.2.1, was on 12 July. Since then there have been 27 releases to npm, ending with 6.0.0 on the morning of 8 October, and three of the major versions came in four days (4.0.0 on 5 October, 5.0.0 on the 7th, 6.0.0 on the 8th). 6.0.0's breaking change was that MCP path inputs must now be absolute. The 5.0.0 changelog section alone has 140 bug-fix entries. GitHub shows 1,467 commits, 23,204 stars and 2,553 forks as I write; the README's "20,000 GitHub stars" banner is already out of date.
Downloads tell a sharper story than stars. The npm API reports 1,693 downloads for 5 September to 4 October, of which 1,291 came in the last week, and 4,048 on 5 October alone. Before 3 October the package was getting 12 to 18 downloads a day.
The engineering discipline is not young. There are 629 test files (280 next to the source, 349 under tests/, 294 of them boundary tests that drive the real MCP server with the pinned client SDK), and a grep for it( and test( blocks finds about 2,700. AGENTS.md bans module mocks in favour of production seams and says real Hopper, Ghidra and browser claims "cannot be replaced by mocks". CI runs on every push and pull request on Ubuntu, macOS and Windows runners; there are separate workflows for real Hopper (macOS and Linux), real Ghidra on Windows, real browsers, real JavaScript recovery and so on. The catch is that the real-Hopper and real-Ghidra-on-Windows lanes run on self-hosted runners and are triggered by hand (workflow_dispatch), so a pull request can merge without them. The docs are candid about that too: one Ghidra recovery flow "was exercised on one real Linux stdio connection [...]. It is not a Windows/macOS coverage claim."
I would describe it as carefully built and not yet stable. Expect the contracts to move under you between weekly releases, and pin the version in your MCP config if you build anything on top of it.
Licences, the law, and what REA says about either
REA itself is MIT. What it drives is not uniformly so. Ghidra is Apache-2.0, free to use and redistribute. IDA is commercial. Hopper is commercial with a free demo whose limits, according to Hopper's download page, are no saving or exporting of disassembly, a disabled debugger, and sessions capped at 30 minutes. REA's installation guide says Hopper "is separate commercial software with its own license", and its setup asks before installing it. On Linux, though, REA's unattended path is built around the demo: it installs the vendor's demo package, verifies its checksum, and clicks the demo button for you on a hidden display, once for every analysis session, each of which is a fresh Hopper process. I could not find anything in Hopper's terms on the download page about automated use of the demo, and I have not read the full licence agreement; if you plan to run this in CI, that is a question for Hopper's vendor, and buying a licence answers it.
On the law of reverse engineering itself, REA says one sentence, at the bottom of the README: it "provides tools for lawful reverse-engineering research, analysis, and reconstruction. You are responsible for obtaining any required authorization and complying with applicable laws." I found nothing in the docs, the skill or the website about DRM, anti-tamper, or terms of service, and the skill has no instruction to stop at a protection mechanism.
The ground rules, as I understand them, go roughly like this. I am not a lawyer, and this varies by country. In the US, courts have treated disassembly to understand an unprotected functional element as fair use when it was the only way to get at it (Sega v. Accolade, 1992), and DMCA §1201(f) allows circumventing a technical protection measure only to achieve interoperability of an independently created program, and only to the extent needed. A licence agreement that forbids reverse engineering can still bind you there (Bowers v. Baystate, 2003). In the EU, Article 6 of the Software Directive (2009/24/EC) permits decompilation by a lawful user when it is indispensable for interoperability, limited to the parts needed, and not used to build a substantially similar program; Article 5(3) lets you observe and test a program to learn the ideas behind it; and contract terms that override those rights are void under Article 8. Breaking DRM or anti-tamper for any other purpose is a different matter everywhere, and this article does not cover how.
The distinction that matters most for REA's pitch is between learning an idea and copying an expression. "Understand how Notes does search and build a similar feature" is the first: ideas, algorithms and behaviour are not protected by copyright, and writing your own implementation after studying them is how most software is made. Copying the decompiled code, the assets or the strings into your product is the second. The DX-Ball project sits in between in an interesting way: a byte-exact reconstruction is by design a re-expression of the original's code, which is why it publishes its own C under MIT but ships none of the game's binaries, art or audio and asks you to supply your own copy. Projects like that have existed for years in a grey zone; the interoperability exceptions were not written with them in mind.
Would I use it
For JavaScript and Electron apps, yes, today, with the caveat above: run it, read the graph, and when the counts come back empty on minified code, open the files rather than believing the zero. It installs in seconds with no native dependencies.
For native work, I would use it with Ghidra rather than Hopper, partly for the licence and partly because the Ghidra bridge is the more complete of the two. What I would actually be buying is not decompilation, which I already have, but the dossier: pseudocode and assembly together, callers and data references in one call, and an Evidence id I can quote back to the agent when its explanation sounds too sure.
For me, the most useful idea in the project is not any one tool. It is the expectation, written into the code, that every answer arrives with the evidence and the limitations attached. Decompilers have never worked that way, and agents built on them certainly have not. Whether models learn to respect the receipts is a separate question, and REA cannot answer it for them.
Related on this site: the GTA V browser port teardown, which is what reverse engineering looks like when you have only the shipped files and a lot of patience, and OpenDLSS-NR, a reimplementation whose author says it was reverse engineered but does not show how.
How I checked
I shallow-cloned morluto/rea at commit 84a17d5 (9 October 2026, +08:00) and read README.md, AGENTS.md, SECURITY.md, docs/ (installation, MCP contracts and prompts, the JavaScript, managed-code and browser guides), the skill in skill-src/, both bridges, src/domain/evidence.ts, the Electron static analysis code, the Hopper and Ghidra launchers, the Windows addon's README, the workflows under .github/ and the changelog. Line numbers in the quotes are from that commit. The tool count and every contract in the catalog widget come from the published npm package rea-agents@6.0.0, installed with --ignore-scripts into a scratch directory, by importing TOOL_CONTRACTS and converting each schema with Zod's toJSONSchema. Test-file and line counts are find and wc over the clone; the test-block count is a grep, so treat it as approximate. Stars, forks, creation date, commit count and contributors come from the GitHub REST API, and downloads and release dates from the npm registry and downloads API, on 8 October.
The DX-Ball walkthrough quotes REA's case-study page (byte-identical in the repository and on the live site) and its evidence note, website/evidence/dx-ball-sound-pan.md; the screenshots are of the live page. I shallow-cloned the DX-Ball reconstruction repository to read its README and src/gameplay.c, but did not build or run it, and I did not have DX-Ball, Hopper, Ghidra or IDA, so nothing native was re-run. The GCC comparison is my own compile of an equivalent C function with GCC 13.3.0. The Electron results are my runs of the 6.0.0 CLI on an app I wrote, in four spellings, the minified one produced with esbuild 0.28.2; nothing third-party was analysed. The Hopper demo limits are from hopperapp.com's download page.