# treg: a people-search router priced by expected cost per hit

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/treg
> date: 2026-10-06
> tags: agents, tool-calling, benchmarks, licensing

Jason Zhou of Superdesign posted a 40-second film on X calling treg the
"OpenSource Clay killer": **85% cheaper than Clay**, **most accurate on
people-search bench**, **25x faster**, and you run it from Claude Code or Codex.
The repository came in the next post. I cloned it, read the routing code and the
catalog, pulled the frames out of the film, and checked each number against
whatever primary source exists for it.

<RepoCard repo="superdesigndev/treg" note="Read at f93ae9b (2026-10-06). Python package tools-registry 0.22.0; 4,090 tracked files, about 90,000 lines of Python under src/treg." />

The short version: the thing that does the work is a small, well-reasoned
router, about 1,000 lines across `plan.py` and `route.py`, sitting on a large
hand-curated price catalog. The router deserves the attention. The launch
numbers mostly cannot be checked, and the one chart that can be checked is
labelled wrong.

<Figure
  src="https://ai.thesatyajit.com/articles/treg/fig1.png"
  alt="treg's README banner: a dark treg badge, the headline 'OpenRouter for agent tools', the line 'One unified key for 2000+ tools, priced per call, open source', a masked trg_live_ token, and logo pills for Semrush, Moz, Majestic, SerpApi, DataForSEO, SE Ranking, TikTok, Instagram, YouTube, X, Reddit, LinkedIn, Hunter, Lusha, PDL, Apollo, Crunchbase, Coresignal, Google Ads, Meta Ads, GA4, Pinterest and Snapchat, ending in '+2,000 endpoints'."
  caption="How treg describes itself: one token in front of thousands of paid endpoints (the project's README, docs/assets/treg-hero.png)."
/>

| | |
|---|---|
| Repository | [superdesigndev/treg](https://github.com/superdesigndev/treg), Python 3.12+, FastAPI server and a `treg` CLI |
| Commit read | `f93ae9b`, 2026-10-06 |
| Catalog | 3,851 endpoints in 109 provider files under `src/treg/catalog/` (measured) |
| Routing | 90 capability contracts, 449 adapters (measured) |
| Licence | Apache 2.0 plus "Additional Terms": no hosted, managed or embedded service for third parties without written permission |
| Hosted | [treg.to](https://treg.to/people-search), prepaid balance, \$1.00 free for a new team |

## What treg is

treg is not a people database. It holds no contacts. It is three things stacked
on each other.

**A catalog.** One YAML file per provider (Hunter, Tomba, Apollo, Lusha, Exa,
DataForSEO, TikHub and 100 more), each endpoint carrying its parameters, an
example response captured live, and a `cost` block. The cost block is the
interesting part: a type (`per_success`, `per_call`, `per_result`, `free`), a
value in the provider's own unit, and a note on how that was established. Most
of them read like lab notes. Tomba's says a repeat of the same query within a
month bills zero, "confirmed live" six minutes apart; Lusha's says the docs claim
search is free and the billing field says one credit per 25 rows. Credits become
dollars through `fx.yaml`, which divides each provider's cheapest public plan by
its credits (Tomba: \$89 for 10,000 credits, \$0.0089 a credit). That makes
these list-price upper bounds, as the file itself says.

**A proxy that injects credentials.** You build the real upstream request and
prefix it with `https://treg.to/call/`. treg resolves the tool by host, adds the
key server-side, and relays everything else byte for byte, stripping only
hop-by-hop headers, its own `X-Treg-*` headers and its session cookie. Whose key
gets used follows a fixed ladder (README, "the credential ladder"):

1. a tool your team registered for that provider;
2. a secret your team stored for it;
3. a verified public route that needs no key, which is free;
4. otherwise treg's own key, billed to the team's prepaid balance.

Your own key always wins, and calls on it are never metered.

**A meter.** A platform-key call reserves its estimated cost from the balance
and settles afterwards from what the provider actually reported. Statuses that
mean "our key was rejected or throttled" (401, 402, 403, 429 and friends) are
never charged. An endpoint with no published price is refused, not served
free. The margin over the provider's rate is a deploy setting,
`platform_margin`, whose default in `config.py` is `0.0`; that matches the
site's "0% markup", but what the hosted deployment sets is not in the repo.

That is the general product: "OpenRouter for tools". People search is one
aisle of it, and the only aisle the launch talks about.

<Figure
  src="https://ai.thesatyajit.com/articles/treg/fig2.png"
  alt="The top of treg.to/people-search: the headline 'Hermes for people search', a line saying one skill gives an agent 1B+ contacts across Apollo, Hunter, Tomba and People Data Labs with 109 providers behind one token, a link reading '27-person receipt: 20 verified emails for $0.58', and a demo table of companies (Modal, Baseten, Fireworks, Together, Modular) with a side panel labelled 'treg · people.email.find, cheapest first, misses free' showing Tomba."
  caption="The landing page the post links to, rendered by me in headless Chromium on 2026-10-06. The 109-provider count it prints matches the catalog (treg.to/people-search)."
/>

## How a routed people search works

Calling one provider is easy. The hard part of "find this person's work email"
is that 23 providers can answer it, each wants a different request shape, each
bills differently, and each misses on different people. Clay's waterfall column
is a hand-ordered list of them. treg's answer is a routed endpoint,
`POST /call/treg.people.email.find`, built from two data files.

**A contract** (`catalog/contracts.yaml`) says what the question is, in
provider-free terms. For `people.email.find` it lists four identity variants:
`{full_name, domain}`, `{first_name, last_name, domain}`, `{linkedin_url}` and
`{linkedin_handle}`, plus `derive` rules so a caller may send whichever it holds
(`first_name` is `split_first(full_name)`, a handle is pulled from a URL). It
names the output core every answer must fill (`email`, required; `confidence`,
`verified`) and what a miss looks like: `email == null`.

**An adapter** per endpoint (`catalog/adapters.yaml`) maps the contract onto
that provider: which variants it `accepts`, where each field goes
(`full_name` → `body.data.full_name` for Prospeo), where the answer comes back
(`person.email.email`), and its own miss predicate. Adapters are verified at
catalog load by round-tripping the endpoint's captured test request and example
response; one that fails is never a routing candidate. At this commit 23
adapters route `people.email.find` and 23 route `people.search` (measured).

The router then does four things.

**1. Keep only candidates that can answer.** An adapter that does not accept the
identity you sent is dropped with a reason ("needs `{linkedin_url}`"). For
`people.search` there is a sharper rule: `company_domain` is a *scoping* key, so
a provider that cannot send it is dropped outright. The comment explains why: a
title-only search asked for the CEO of one company "returns CEOs of any company",
and each of those bills as a hit.

**2. Rank them.** `rank()` in `routing/plan.py` sorts by a tuple:

```python
# src/treg/domain/catalog/routing/plan.py, rank() — the sort key
(tier,                      # own tool/credential 0, anonymous route 1, treg's key 2
 pref,                      # X-Treg-Route-Prefer, or the contract's default order
 -covers(variant),          # how many of the caller's keys this variant uses
 len(c.ignored),            # filters the caller sent that this adapter drops
 c.expected_cost_per_hit,   # price x P(billed) / P(hit)
 p50_ms, last_ok_days, endpoint_id)
```

The cost term is the one worth stopping on:

$$
\text{cost per hit} = \frac{\text{price} \times P(\text{billed})}{P(\text{hit})}
$$

Here $P(\text{hit})$ is the endpoint's measured hit rate once it has at least 50
samples, else its success rate, else 1. $P(\text{billed})$ equals
$P(\text{hit})$ when the endpoint is priced `per_success` and 1 otherwise.
So for a provider that charges only when it finds something, the hit rate
cancels and the expected cost per hit is just its price. For one that charges
every call, a 50% hit rate doubles it. Every one of the 13 name-and-domain email
finders in the catalog is `per_success` (measured), which means that for the
commonest request the ranking is pure price, and live hit rates only break ties
through latency. That is a reasonable outcome, and a consequence the code never
states.

The `len(c.ignored)` term sits *ahead* of price on purpose. Its comment records
the incident: a search for `{q, title, location: London, country: GB}` went to
the cheapest child, which mapped neither geo filter, and "returned people in
Bengaluru and San Francisco — reported as a hit, \$0.0025". A provider that would
answer a looser question now ranks last among equals, and the response names
what it ignored in `X-Treg-Ignored-Filters`, or refuses with an unbilled 422 if
you send `X-Treg-Route-Strict-Filters: 1`.

**3. Walk the ranking as a waterfall.** `route.py` calls each candidate through
the same `execute_call` path as any proxied request, so holds, audit rows and
settlement are the ordinary ones. On an error (a 5xx, 429 or 402 from the
vendor) it tries the next, at most two extra. On a miss it continues, unless you
sent `X-Treg-Route-Waterfall: 0`. On a hit it stops. A candidate whose price
would push the running spend past `X-Treg-Route-Max-Cost` is skipped; the
default is \$1.00, "a runaway guard, not a budget". For list answers,
`X-Treg-Route-Min-Results: N` marks a short answer as weak and keeps looking,
bounded at two extra providers, because unbounded it took the benchmark's
lookup set from \$1.76 to \$22.35 over 28 queries "for answers that were already
right".

**4. Say what happened.** The answer carries the contract's core output, the
provider's raw body, and `_treg: {served_by, tried}` with each attempt's outcome
and charge. A hit whose `verified` is not true also carries `_treg.advice`
telling the agent to run `treg.people.email.verify` before sending. The contract
comment says why: one team "bounced on 73 of 79 addresses" that were unverified
directory rows.

The widget replays this with the catalog's own prices. Pick the identity, slide
the rung that actually holds the person's email, and toggle an own key.

<RouteWaterfall />

Three things show up quickly. With a name and a domain, the first five rungs
cost under two cents each and a miss on any of them is free, so finding a person
on rung four costs that rung's price and nothing else. Connecting your own
Hunter key puts Hunter first at zero, whatever it costs Hunter's other
customers. And on the LinkedIn identity the one `per_call` endpoint, HarvestAPI
at \$0.02, slides down the ranking as you lower its hit rate, and bills you on
every miss when the waterfall passes through it.

## Using it

```bash
curl -fsSL https://treg.to/install.sh | sh      # CLI + skills
treg login                                      # GitHub, --email, or --token for agents
treg catalog search "find a work email"         # by the job, not the vendor
treg catalog get treg.people.email.find         # the plan, prices, what each child accepts
treg balance
```

A routed call from an agent, without the CLI:

```bash
curl -s https://treg.to/call/treg.people.email.find \
  -H "X-Treg-Token: $TREG_TOKEN" \
  -H "X-Treg-Route-Max-Cost: 0.10" \
  -H "content-type: application/json" \
  -d '{"full_name": "Alexis Ohanian", "domain": "reddit.com"}'
```

In Claude Code it is `/plugin marketplace add superdesigndev/treg` then
`/plugin install treg@treg`; the skill walks the agent through the CLI and
`treg mcp install`. The skill file (`skills/treg/SKILL.md`, 420 lines) is
mostly guardrails: do not retry a find, every hit bills; verify every address;
a routed search tells you which filters it dropped. The "run GTM in Claude Code
or Codex" part of the launch holds: it is a plugin, a skill and an MCP server.

Self-hosting is `uv sync`, `uv run python -m treg upgrade`, `uv run python -m
treg`, on SQLite by default. A self-hosted registry has no platform keys, so
the catalog becomes a set of adapters for your own keys. That is useful, and
it is the only kind of hosting the licence allows without asking.

## Checking the launch claims

### "85% cheaper than Clay"

The number comes from the film, not from the repository or the landing page.

<Figure
  src="https://ai.thesatyajit.com/articles/treg/fig3.jpg"
  alt="A frame from the launch film titled '85% cheaper than Clay', subtitled 'cost of 100 correct work emails, at each tool's list price'. Bars: treg $0.58 (highlighted, WINNER), Clay $3.77, Apollo $2.91, Deepline $3.40, Hunter $2.63. Footer: 'same 100 leads, name + domain in, graded against the published address · Oct 2026 · full table at treg.to/people-search'."
  caption="The cost claim, at the frame where every counter has settled. These are the publisher's numbers; no run log or table for them is published (launch film in Jason Zhou's X post, 0:09)."
/>

The arithmetic inside the frame holds: 1 − 0.58 / 3.77 = 84.6%, which rounds to
85% (reasoned). The film's closing table gives the same five tools as a cost
per 100 *leads*: treg \$0.53 at 92 correct, Clay "~\$3.35" at 84 correct. Dividing
cost by correct answers reproduces the chart for treg (0.53 / 0.92 = \$0.58),
Apollo and Hunter, but gives \$3.99 for Clay rather than \$3.77 (reasoned). The
tilde is the film's own: Clay's figure is an estimate "at each tool's list
price", not a bill. On either Clay number the saving is about 85%.

What I cannot check is anything underneath. The footer says "full table at
treg.to/people-search". On 2026-10-06 that page has no such table, no list of
the 100 leads, and no mention of Clay's result beyond its \$149/mo plan price
(measured, by fetching and grepping it). The page's own self-playing demo meters
its 100-lead run to \$1.76, a third number for what reads as the same film.

What *is* checkable is the landing page's smaller receipt: "27 people, 20
verified emails, \$0.58". Its CSV is in the repo
(`src/treg/workflow_runs/find-and-verify-a-lead-list.csv`), and it reproduces.
50 companies, 27 kept by the fit gate, 18 emails from Hunter and 3 from Kitt on
Hunter's misses, 6 found nowhere, 20 verified valid and 1 unknown (measured). At
catalog prices that is 18 × \$0.0245 + 3 × \$0.005 = \$0.456 to find (reasoned).
The page adds \$0.125 to verify, which is exactly 20 definitive checks at
LeadMagic's \$0.00625 with the unknown verdict free; the CSV does not name the
verifier, so that match is mine. Total \$0.58, or \$0.029 per usable email
(reasoned). It is a different run from the
film's, and the coincidence of the two \$0.58 figures is worth knowing before
quoting either.

One smaller staleness: the page calls \$0.0089 (Tomba) the "cheapest email-find
on the catalog today". It is hard-coded in the HTML. The catalog at this commit
has Kitt at \$0.005 per found email and QuickEnrich at \$0.004834 (measured).

### "25x faster"

<Figure
  src="https://ai.thesatyajit.com/articles/treg/fig5.jpg"
  alt="A frame from the launch film titled '25x faster than Clay', subtitled 'median time to a found work email'. On the left a faded Clay panel with a timer at 10.9 s and five names still 'finding…'; on the right a highlighted treg panel at 0.44 s listing five names with emails such as maya@baseten.co and sam@dagster.io. Footer: 'same 292 leads, name + domain in · full table at treg.to/people-search'."
  caption="The speed claim. Clay's timer is animated and stops at 11.2 s a few frames later (launch film in Jason Zhou's X post, 0:15)."
/>

Clay's counter stops at 11.2 s; treg's reads 0.44 s. 11.2 / 0.44 = 25.5, so
"25x" is the film's own ratio (reasoned). It is a median over 292 leads, a
different set from the cost chart's 100, and nothing else about the run is
published. The direction is believable from the mechanism: one synchronous API
call to a cheap finder against a spreadsheet tool that queues each row. The
factor is the publisher's.

### "Most accurate on people-search bench"

There are two accuracy claims, and they are not the same test.

The film's is its own: 92 of 100 leads with a known answer, against Clay 84,
Hunter 84, Deepline 80 and Apollo 77 (reported). Like the cost chart, the
leads and grading are not published.

<Figure
  src="https://ai.thesatyajit.com/articles/treg/fig4.jpg"
  alt="A frame from the launch film titled 'Most accurate', subtitled 'correct work emails, 100 leads with a known answer'. Bars: treg 92% (WINNER), Clay 84%, Apollo 77%, Deepline 80%, Hunter 84%."
  caption="The film's accuracy chart: treg's own 100-lead test, not the LessieAI benchmark the landing page cites (launch film in Jason Zhou's X post, 0:12)."
/>

The landing page's is the named benchmark: "People Search Bench by LessieAI —
119 real tasks, % answered correctly. Same agent, with and without treg", with
a headline duel of 43% for Claude Code alone against 78.2% with treg on B2B
prospecting, and four category charts. [PeopleSearchBench](https://arxiv.org/abs/2603.27476)
is real and its repository publishes per-category scores for four platforms.

<Figure
  src="https://ai.thesatyajit.com/articles/treg/fig6.png"
  alt="LessieAI's 'Performance by scenario' chart: horizontal bars per scenario. Influencer/KOL: Lessie 62.3, Claude Code 43.2, Exa 41.6, Juicebox 31.1. Expert/Deterministic: Lessie 70.4, Exa 61.2, Claude Code 57, Juicebox 44.2. B2B Prospecting: Lessie 60.6, Exa 55.2, Juicebox 51.4, Claude Code 43. Recruiting: Lessie 68.2, Juicebox 65.7, Exa 64.7, Claude Code 50.5."
  caption="The benchmark's own chart. treg does not appear in it (LessieAI/people-search-bench README, assets/performance-by-scenario.png)."
/>

I recomputed treg's competitor bars from LessieAI's `summary.json` files. Every
one reproduces as the plain mean of three dimensions: Lessie's B2B 60.6 is
(62.8 + 63.5 + 55.5) / 3, Claude Code's 43.0 is (43.0 + 42.3 + 43.6) / 3, and so
on for all twelve (measured). That settles what the chart is, and it is not "%
answered correctly". The three dimensions are padded nDCG@10 over web-verified
relevance grades, an effective-coverage score, and an information-utility
score. A 43 does not mean Claude Code answered 43% of B2B tasks correctly.

<BenchDecompose />

Three more things the comparison leaves out:

- **Juicebox is dropped.** It is the fourth platform in the benchmark; its
  recruiting score of 65.7 sits between Lessie and Exa (measured). Leaving it
  off does not change treg's rank, but the chart is not the benchmark's field.
- **treg's own numbers have no source.** 80.0, 78.2, 76.3 and 62.9 are
  hard-coded in `people-search.html` and `routers/web.py`. LessieAI's repository
  lists no treg submission, and treg's repository has no run log, per-query
  output or per-dimension breakdown. The closest thing is a design note in
  `docs/context/architecture/catalog.md`: on the 30 recruiting briefs, a filter
  fix moved treg's routed search from 55.4 to 69.2 "overall" on 2026-08-29,
  "ahead of the published Lessie 68.2". The page says 80.0. Either a later run
  exists and is not in the repo, or the page is not that run.
- **"Same agent, with and without" is two different runs.** The "Claude Code
  alone" bar is LessieAI's Claude Code run from its paper. The "with treg" bar is
  Superdesign's own. Same benchmark, not the same harness on the same day.

The repository is candid about part of this. The Enrich Arena's benchmark
page calls its charts "a published snapshot, not a new benchmark run", and its
design doc says "the original repository's current figures differ from the
landing" and that the page "does not manufacture an overall score". The landing page's subtitle did not
get the same care.

### "Open source"

The licence is Apache 2.0 with Additional Terms that "take precedence over the
Apache License to the extent of any conflict". The first term forbids using the
software "to provide a hosted, managed, or embedded service to third parties"
without written authorisation. That is a field-of-use restriction, which the
Open Source Definition does not allow, so this is source-available in the usual
sense of the words. For a team that wants to run its own registry, nothing is
lost: self-hosting for your own organisation is "expressly permitted and
encouraged", and calling the hosted API from your own product is allowed.

## What the launch leaves out

Superdesign's own marketing draft, `marketing/rebuild/01-rebuild-clay.md`, is a
better review of the Clay comparison than the film. It lists where Clay is
still the right buy: you want a spreadsheet, not a script; you need a research
column that reads each company's website; you run fifty tables unattended and
want Clay's queueing and rate-limit handling: "Here, that's your job." It
also says plainly that Clay's moat is "150-plus pre-negotiated data providers".
treg's answer to that moat is its catalog: real, long, and priced at list rates
it documents per provider.

A reply under the launch post made the remaining point: a people-search score
is the easy part, and the test is how stale the data is after 90 days and how
many sends bounce. treg holds no data of its own, so staleness is whatever the
answering provider's is. Its defence is the verify step it keeps pushing, a
fraction of a cent per address, which caught the bounces in its own incident
notes.

## Verdict

| Claim | Status |
|---|---|
| Run it from Claude Code or Codex | Holds: plugin, skill and MCP server in the repo |
| 85% cheaper than Clay | Arithmetic consistent within the film; underlying run not published |
| 25x faster | 11.2 s / 0.44 s in the film; underlying run not published |
| Most accurate on People Search Bench | Competitor bars reproduce; the label is wrong, Juicebox is omitted, and treg's own scores have no published run |
| 27 people, 20 verified emails, \$0.58 | Reproduces from the committed CSV and catalog prices |
| Open source | Apache 2.0 with a hosted-service restriction: source-available |

The piece I would take from treg is not a number from the film. It is the
router: contracts that state the question without naming a vendor, adapters
that are verified against captured fixtures before they may serve, a ranking
that puts "answers the question you asked" ahead of "is cheapest", and a cost
term that is honest about what a miss costs. If you build an agent that calls
paid APIs, `plan.py` is 183 lines and worth reading whole. For a sibling story
about benchmark multipliers that come apart into simpler factors, see
[the WindTunnel decomposition](/articles/webmcp-windtunnel); for another way to
keep an agent's paid side effects on a leash, [Cloudflare OS's
gatekeepers](/articles/cloudflare-os); and for harnesses read for where they
check the agent's work, [three agent harnesses](/articles/agent-tools-week).
