# OpenQodex: thirteen scanners, and a reviewer that cannot skip a finding

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/openqodex-code-review
> date: 2026-10-07
> tags: agents, agentic-coding, developer-tools, security, harness

Two numbers in the launch post made me clone this one. Siddhant Mohan's [announcement](https://x.com/siddhantmohan1/status/2107645119130440006) says the review "you pay \$60 a developer for" is now \$0, and that it is "100+ scanners plus an AI reviewer that has to answer every finding, on your own Claude Code or Codex login, before you push." One of the replies asked the question I had: "Who is paying developers \$60 for a review? … What differentiates yours other than 'big number'?"

So I read the repository, [openqodex/openqodex](https://github.com/openqodex/openqodex) at commit `549d330` (version 0.8.1, 2026-10-06): about 25,600 lines of TypeScript outside the tests and 19,200 lines of tests. I expected a prompt wrapped around semgrep. That is not what is there. The engineering is careful, and some of it is worth copying. The pitch is a different matter. The big number is wrong, and "before you push" means something weaker than it sounds.

## Counting the scanners

The launch video has a slide for the number.

<Figure
  src="https://ai.thesatyajit.com/articles/openqodex-code-review/fig2.png"
  alt="A slide reading '100+ scanners' over a grid of thirteen tiles: gitleaks, semgrep, osv-scanner, trivy, actionlint, shellcheck, hadolint, brakeman, bandit, eslint, ruff, sqllint, checkov, and a fourteenth tile reading 'any scanner by its GitHub link'. Under it: 'plus an AI reviewer that has to answer every finding'."
  caption="The '100+ scanners' slide. Three of the thirteen named tiles (trivy, eslint, checkov) are not built into OpenQodex, and three built-ins (oxlint, golangci-lint, rubocop) are missing from it (a frame of the launch video, about 13 s in)."
/>

The code does not say 100 anywhere. It says thirteen, in a comment on the list itself, `packages/scanners/src/adapters/index.ts:1`: "The thirteen builtin scanners." The README says "Thirteen built-in scanners." The product page at qodex.ai says "13 built in." I counted the `ADAPTERS` array (`index.ts:42-67`) and got the same thirteen: semgrep, gitleaks, sqllint, osv-scanner, actionlint, hadolint, shellcheck, ruff, brakeman, rubocop, bandit, oxlint and golangci-lint. One of them, sqllint, is not a downloaded tool at all. It is three Postgres migration rules written inside OpenQodex.

The slide's grid is not that list. Trivy shows up in the repo only as the README's example of a custom scanner, checkov only as a test fixture for the custom-scanner downloader, and eslint not at all: the JavaScript linter is oxlint. A grep for "100+", "hundred" and "\$60" over the whole repository finds one hit, a comment about a symbol with "hundreds of callers" in the graph tests.

There are two generous ways to reach a hundred, and I tried both. The first counts the last tile, "any scanner by its GitHub link", as unbounded. That is a real feature (`openqodex trust` shows you the release asset, its sha256 and the run line, and asks yes or no), but a feature that lets you add scanners is not a hundred scanners. The second counts rules instead of tools. semgrep runs three registry packs, `p/default`, `p/security-audit` and `p/secrets`. I fetched all three on 2026-10-07 and got 1,073, 225 and 52 rules, 1,089 distinct. So "1,000+ rules" would have been true, and ruff, shellcheck and hadolint would add hundreds more. "100+ scanners" is not.

There is also a second catalogue that the slide does not mention: 48 "lenses" in `packages/core/lenses/`, each a markdown file describing one bug pattern (a floating promise, an open redirect, `SECURITY DEFINER` without a `search_path`) with a file glob, a regex over the changed lines, and a confidence floor. A change that trips a lens gets its text in the reviewer's brief, at most four per brief (`MAX_LENSES = 4` in `lenses.ts`). These are prompts, not scanners, and the code is careful to call them "patterns to weigh."

## What the scanners actually see

Here is the part the count obscures, and the part I like. OpenQodex does not point thirteen tools at your repository. It points them at your change.

`openqodex review` starts by working out the change: the commits not yet pushed plus everything uncommitted, untracked files included. It copies that state into a frozen snapshot under `~/.openqodex/checkouts/`, and everything after this point reads the snapshot, never your working folder. A hash of every snapshot file is taken before the reviewer starts and again after it ends; if they differ, the review is incomplete.

Then each adapter answers one question, `wants(changedPaths)`: does this change hold a file I read? semgrep and gitleaks say yes to anything. osv-scanner wants a lockfile, hadolint a Dockerfile, actionlint a workflow under `.github/workflows/`, brakeman a Ruby file in a repo that has both a `Gemfile` and an `app/` folder. Only the scanners that say yes get resolved, which matters because resolution can mean a download: each pinned tool lands in `~/.openqodex/tools/<scanner>/<version>/` on first use, checked against a sha256 in the package. An install that runs past 45 seconds keeps going in the background, and the report lists that scanner as `installing` rather than holding the review hostage.

The ones that want the change run in parallel, `Promise.all` over the builtins and any approved custom scanners (`run.ts:95-101`). Every run is wrapped in a `guard` that turns a throw into a status. The file's header says it plainly: "Static analysis is additive context, never a gate on its own." A scanner that is missing, broken or offline shows up in the report as `not_installed`, `failed` or `disabled` with a one-line reason, and never changes the exit code.

What comes back goes through one pipeline, in this order:

1. Paths are rebased onto the repository root. This is more delicate than it sounds. ruff reports absolute paths resolved through symlinks, and on macOS `/var/folders` is a link to `/private/var/folders`, so a naive prefix match silently drops every finding. The comment above `toRunDirRelative` explains the bug it exists to prevent.
2. Hits are cut to the changed lines. A finding survives only if its line range touches a line the change added or modified (`filter.ts`). Pre-existing problems in code you did not touch are gone.
3. Hits in fixtures, mocks, snapshots and `testdata/` folders are dropped, and counted. Test files themselves are kept on purpose, because "a hardcoded admin token in a test runner" is worth surfacing.
4. Hits from two scanners on the same span, in the same coarse class, merge. The classes are `secret`, `injection` and `auth`, matched by regex on the rule id (`ruleClassFor`); anything else only merges with exact repeats of itself. gitleaks and semgrep reporting one Stripe key on one line become one candidate, and ties go to whichever scanner comes first in `ADAPTERS`, which is why semgrep is listed before gitleaks.
5. What is left is sorted by severity and numbered `c1`, `c2`, and so on. Those ids are what the reviewer has to answer.

Two kinds of candidate do not come from any scanner. A change that adds a suppression comment, `# nosec`, `// eslint-disable-next-line`, `gitleaks:allow` and the rest, gets a candidate on that line from the scanner it silences, because that scanner will now report nothing there. A change that edits a scanner's settings or ignore file (`.gitleaks.toml`, `.semgrepignore`, a `[tool.ruff` table in `pyproject.toml`) gets one too. This is the cleverest small thing in the repo. The scanner obeys the comment, so its own output can never show what the comment hid; only something reading the diff can. The per-scanner comment readers in `suppression.ts` go as far as matching shellcheck's heredoc rules and BuildKit's `RUN <<EOF` parsing, and `docs/scanners.md` lists each marker with the scanner source file it was checked against.

The repository ships a demo that makes all of this concrete: a small Flask shop with twelve bugs planted on purpose, listed with their expected detectors in `examples/demo-repo/expected.json`. The widget below takes that diff through every stage. The scanner gating, the merges, the rejection messages and the verdict logic come from the code. The reviewer's answers are my illustration of the format, not output from a run.

<ReviewPipeline />

Run on the demo, eight scanners want the change and five have nothing to check. semgrep and gitleaks take everything; ruff and bandit take the three Python files; osv-scanner, actionlint, hadolint and shellcheck each take one file. sqllint, brakeman, rubocop, oxlint and golangci-lint sit out. That matches the line in the launch video exactly, "Scanners: 8 ran, 5 had nothing to check, 26 candidates to check."

Two of the twelve planted bugs have no scanner at all, and they are the interesting ones. `app/server.py` changes the page offset from `(page - 1) * PAGE_SIZE` to `page * PAGE_SIZE`, so page one skips the first page of results. The `Dockerfile` loses its `USER 10001` line, so the container runs as root again. semgrep does notice the missing user, but it reports it on line 11, the `CMD` line, which the change did not touch. The changed-line filter drops it. The demo's own notes say "Only a reviewer reading the diff finds it." That is the case for the second half of the tool.

## "Has to answer every finding" is a validator

I expected the phrase to mean a sentence in a prompt. It means a function, `checkSubmission` in `packages/core/src/finalize.ts`, and it is strict.

The reviewer's whole answer is one JSON object: a summary, a list of `findings`, and a list of `dropped` candidates. A finding that raises a candidate names its id and must set `source` to that candidate's token (`semgrep:python.lang.security...`); the script rejects a mismatch. A drop needs a reason and a file and line "that show why", and the cited line has to exist in the snapshot. Then comes the loop that gives the feature its name:

```ts
// packages/core/src/finalize.ts:519-523
for (const c of live) {
  const n = dispositions.get(c.id) ?? 0;
  if (n === 0) errors.push(`candidate ${c.id} (${c.filePath}:${c.lineStart}) has no disposition: raise it in a finding or add it to dropped with a reason and a line`);
  if (n > 1) errors.push(`candidate ${c.id} has more than one disposition: raise it once or drop it once`);
}
```

Exactly one disposition per candidate, no more and no less. Findings must start on a line the change added or modified, or on a line next to a deletion, and span at most 200 lines, because a range of `1` to `999999` would otherwise overlap every change.

There is a second set of rules that I did not expect, and they are about prose. Every `problem`, `consequence` and `fix` is at most two sentences of at most 20 words each, one line, no em dash, and it may not name a scanner or a rule id, because "the report shows the source on its own line." A drop reason gets the same sentence limits. Somebody at Qodex has read a lot of model-written review comments.

Failures are not fatal on the first try. The script collects every broken rule, numbers them, and sends them back into the same reviewer session: "Your answer failed these checks. Fix every one." Then "answer again with the whole JSON object and nothing else." `MAX_CORRECTIONS = 2` (`review-run.ts:69`), so the reviewer gets the brief plus at most two correction rounds. If the answer still fails, the run records "the reviewer's answer still failed N checks after the correction rounds", prints "Review incomplete: this is not a review of the change", and exits 2.

The same correction rounds carry a second guarantee: every changed line has to have been in front of the reviewer. The brief holds the diff, up to 200 KB of it. For Claude Code, a changed range also counts as seen when the reviewer's own `Read` calls covered every line of it, which OpenQodex reads off the event stream. Anything still unseen after an answer is pasted into the next correction message by the tool itself, with three lines of context, up to 4,000 lines a round: "These changed lines were not in front of you yet." Coverage never depends on the model deciding to open a file.

Then it is worth being exact about what the scripts cannot check, because the launch phrasing implies more. They check that every candidate got exactly one answer, that each answer is well-formed, and that cited lines exist. They cannot check that an answer is right. A drop reason of "Same problem as the finding on this line" with any real line number passes every check. A finding raised at confidence 0.6 counts as a disposition, then lands in a "Below the confidence floor (not counted)" list instead of the findings, so it never touches the verdict. That is the documented design ("A wrong finding is worse than a missed one"), and I think it is the right call. But "has to answer" means the reviewer cannot be silent about a hit. It does not mean the hit was judged well. The reply that asked what happens "when the reviewer and the scanner agree… and both are wrong" has no answer in the code, and the README does not pretend otherwise: "No review finds everything. The promise is that every stage runs, every scanner finding is checked, every changed line is put in front of the reviewer, and anything skipped is named."

What does a finished review look like? The launch video shows a run of the demo.

<Figure
  src="https://ai.thesatyajit.com/articles/openqodex-code-review/fig3.png"
  alt="A terminal. 'npx openqodex init' reports OpenQodex set up for Claude Code, Cursor and Codex CLI. 'openqodex review' prints: Reviewing the change against HEAD: 8 files, +38 -14. Scanners: 8 ran, 5 had nothing to check, 26 candidates to check. Reviewer: claude 2.1.292 started. Passed with warnings: 10 findings (2 critical, 4 major, 4 minor). Findings: 1. Critical security: Search query built from request input, app/search.py:14. 2. Critical security: Stripe secret key committed to source, app/config.py:2. 3. Major bug: Pagination offset skips the first page, app/server.py:23, Source: the reviewer."
  caption="The demo run from the launch video: 26 candidates in, 10 findings out, the pagination bug found by the reviewer with no scanner behind it. Note the verdict line: two critical findings and it still says 'Passed' (a frame of the launch video, about 29 s in)."
/>

Twenty-six candidates went in and ten findings came out, at least one of them the reviewer's own (the pagination bug, "Source: the reviewer"). So most candidates were dropped with a reason or merged into a finding that the report counts once. The video does not show the rest of the report, and I did not run the tool, so I cannot say which. The verdict line is the part I would have changed before recording: "Passed with warnings" over two critical findings. I come back to why below.

## Driving your Claude Code or Codex login

"On your own login" is literal. The reviewer is your installed `claude` or `codex` binary, started as a child process without a shell and logged in as you. No API key is involved, and a review spends your plan. What makes it a separate reviewer, rather than the agent that wrote the code grading its own homework, is how much of that binary gets switched off.

For Claude Code, the command line in `packages/cli/src/reviewers/claude.ts:19-53` is:

```bash
claude -p --output-format stream-json --verbose --input-format stream-json \
  --tools Read,Grep,Glob,WebSearch,WebFetch \
  --permission-mode dontAsk \
  --setting-sources "" \
  --settings '{"autoMemoryEnabled":false,"hooks":{},"disableAllHooks":true}' \
  --strict-mcp-config --mcp-config '{"mcpServers":{}}' \
  --disable-slash-commands --no-session-persistence \
  --allowedTools WebSearch,WebFetch
```

Read, search and list, and nothing else: no shell, no edits, no subagent, no MCP server, none of your settings, hooks, memory or `CLAUDE.md`, and none of the repository's. The working directory is the snapshot, and `dontAsk` refuses any read outside it without a prompt. The environment is rebuilt from an allowlist (`PATH`, `HOME`, `USER`, locale, proxy settings and the `ANTHROPIC_*` variables), so `GITHUB_TOKEN` and `NPM_TOKEN` never reach the child. Because the input format is stream-json, the session stays open, and the correction rounds go to the same conversation.

OpenQodex does not trust any of this. It reads the event stream and fails the run if the `init` event lists a tool beyond the allowed ones, any MCP server or a memory path, or if any hook event appears at all. Every tool call is checked from the moment it is requested, whether or not it got a result: each path-bearing input is resolved against the snapshot through the real path of its deepest existing folder, and an attempt outside, even one Claude Code refused, makes the review incomplete. `docs/internal-reviewer-drivers.md` records the canary runs behind each flag, down to a `SessionStart` hook that a terminal wrapper injected and that only `disableAllHooks` stopped. This is the same instinct as [DeepSeek's harness](/articles/deepseek-harness): once a model is in the loop, the rule a human would have applied becomes a script, or it stops being applied.

Codex gets the same treatment with different tools, and with two limits written down. The driver runs `codex exec --ephemeral` with a custom permission profile that confines reads to the snapshot and the system folders, no write entry and no network entry. Because those are config keys a newer Codex could rename or ignore, every review first runs a probe under `codex sandbox` with the same keys: it reads a random canary file outside the snapshot, reads one inside, tries to write, and prints a random marker last. The reviewer starts only if the outside read and the write failed and the inside read worked. The limits: Codex still loads your global `~/.codex/AGENTS.md` (no switch removes it without moving your login), and its event stream does not show every command, since most models run the shell inside a code tool. So the Codex driver is marked `traced: false`, its reads never count toward coverage, and only the brief and the correction rounds can cover a changed line. A Codex review is weaker evidence than a Claude Code one, and the report says so ("Files read: not recorded by Codex").

One default I would change on day one. The web tools are on. The reviewer reads untrusted text (the diff, the code, scanner messages) and private code, and it can fetch URLs, which is the textbook shape for prompt-injection exfiltration. The docs say exactly this, "can be talked into putting that code into a web address", and give the switch, `reviewer_web: off` in `~/.openqodex/config.yaml`. I would ship with it off. The GitHub Action already turns it off.

## What runs before you push

Less than "before you push" suggests. Nothing reviews at push time. There are three hooks, and two of them only look up a record.

The Claude Code and Codex installs add a `PreToolUse` hook on the Bash tool that fires on `git push`. An optional git `pre-push` hook (`openqodex hook install`) does the same for every push, with the exact commit ranges git hands it. Both look for the receipt that `openqodex review` writes at the end of a run, keyed by the change id, in your home folder rather than the repository's `.openqodex/` folder, where a branch could plant one. Neither scans, and neither starts a review. The third hook, for the pre-commit framework, runs the scanners only (`openqodex scan`), and the README says so: "It is not a review."

The decision is one function, `checkPush` in `packages/core/src/push-gate.ts`:

<PushGate />

Three things in it matter. The gate can abstain or deny; it never approves, because an allow would skip your agent's own permission prompt for the push. `review.block_on_severity` is unset by default (`config.ts:28`), and with no threshold nothing ever denies: a missing review is one line of advice, and a finished review is silent whatever it found. That is why the video's demo says "Passed with warnings" over two critical findings. And an incomplete review never blocks, even with a threshold set. That last one is deliberate (a reviewer timeout should not hold your push hostage), but it means a review that failed its own checks is treated more leniently than no review at all, which under a threshold is denied. `OPENQODEX_SKIP=1` lets any push through and says so.

None of this is hidden. The README's own list is accurate: "They warn by default and block only when `.openqodex/config.yaml` sets `review.block_on_severity`." My point is narrower. "An AI reviewer that has to answer every finding … before you push" describes what you get after you set a threshold, not what installs.

## The \$60 and the \$0

<Figure
  src="https://ai.thesatyajit.com/articles/openqodex-code-review/fig1.png"
  alt="A slide reading '0 dollars per developer', with two bars: OpenQodex at 0 dollars and CodeRabbit at 60 dollars, labelled 'per developer per month'."
  caption="The pricing comparison from the launch. The 60 dollars is CodeRabbit's Team plan billed monthly; its entry plan is 30 dollars, or 24 billed annually (a frame of the launch video, about 8 s in)."
/>

The \$60 is real but chosen. CodeRabbit's pricing page lists Essentials, formerly Pro, at \$30 per developer per month, or \$24 billed annually, and Team, formerly Pro Plus, at \$60, or \$48 annually. The launch compares against the second tier.

The \$0 is real in the narrow sense the README states: "A review uses your own Claude Code or Codex plan." The Codex driver's notes measured a real six-line review at 50,384 input tokens and 641 output, in 31 seconds. A one-to-three-minute review on every push comes out of the same quota you code with. The GitHub Action is different again: it runs the full review only when the workflow gives it an `ANTHROPIC_API_KEY`, and "Each review spends the repository's own API credit." Without a key it runs the scanners only.

The product comparison is the more useful one, and the author gave it himself in a reply: CodeRabbit and Greptile "review the PR after it opens, as a hosted service. openqodex runs before you push, on your own claude code or codex login." Qodex also sells a hosted reviewer, and the product page sets the two side by side; the hosted one uses "Two frontier models from different labs" and remembers dismissed findings, and the open one does neither. That last row answers the reply asking whether the same model should review its own code. The reviewer is a fresh process with no memory of the session that wrote the code, but if you wrote the code with Claude Code, by default it is Claude reviewing Claude (`auto` picks the agent you are running in first).

A reply also pointed at [ReviewRouter](https://github.com/777genius/review-router-ai). I cloned it to compare. It is a different shape: a GitHub App plus a CI action, reviewing pull requests in your runner with your own Codex, Claude Code or OpenRouter credentials, with a dashboard, self-hosting, and the option to run several reviewers and require agreement before an inline comment is posted. That agreement threshold is a real answer to the same-model question, and OpenQodex has nothing like it. It has no scanner layer and no forced disposition. The clone I took (`1f6ce29`, 2026-10-07) has no licence file, so it is not open source in the sense OpenQodex is.

## What I would use it for

The engineering here is better than the launch. Scanners gated by the change, cut to the changed lines, merged across tools, and a reviewer that must account for each hit by id, in a process with its tools stripped and its every read checked: that is a sound architecture for pre-push review, and the suppression-comment candidates are an idea I have not seen elsewhere. The docs are unusually honest. Every limit I found, I found written down.

What I would change is the defaults. I would set `block_on_severity: major` in `.openqodex/config.yaml`, set `reviewer_web: off`, and install the git `pre-push` hook rather than relying on the agent hook, which the source itself calls "a reminder about the developer's current work, not the boundary." With those three lines it is the tool the launch post describes. Without them, it is a very well-built report that you are free to ignore.

It sits next to two other things on this site about the same problem: [gdp-ts](/articles/gdp-ts), which moves one class of authorization bug from review into the type checker, and [OpenShell](/articles/openshell), which keeps an agent's boundary in the kernel rather than in its prompt. OpenQodex does the second thing for the reviewer and leaves the first to its scanners.

## How I checked

I shallow-cloned `openqodex/openqodex` at `549d330` and read the CLI, core, scanners and graph packages, the docs and the demo, but did not run any of it; nothing above comes from my own run. The scanner count is the `ADAPTERS` array and the README table. The semgrep rule counts are the three registry packs as served on 2026-10-07; they change over time. The demo's gating, merges and filtered hit come from each adapter's `wants`, `ruleClassFor`, the deletion anchors in `change.ts` and `expected.json`, cross-checked against the "8 ran, 5 had nothing to check" line in the launch video. The figures are frames of the launch video, downloaded through the fxtwitter mirror. CodeRabbit's prices are from its pricing page on 2026-10-07. The reviewer's dispositions in the first widget are illustrative; the rejection messages in it are the strings the code builds. I did not verify that the isolation flags behave as the driver notes record with current Claude Code or Codex releases; those notes cite Claude Code 2.1.289 and codex-cli 0.160.0.
