~/satyajit

treg: a people-search router priced by expected cost per hit

mdjsonmcp

2026-10-06 · 18 min · agents · tool-calling · benchmarks · licensing

Why read this

Hightop 30%

Reads treg's router and recomputes its launch benchmark: the routing is sound, the accuracy chart is mislabelled and the licence bars hosting.

  • Original, source-checked analysis
  • Concrete numbers to act on
  • Open code or weights

Developer tools & infraAPI onlySource-availablePractitioner tool

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
3 of 3: Reproduces the headline result, or shows from primary files it is wrong
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
1 of 3: Relevant for months
Does it affect many?
1 of 3: A specialist community
Only here?
3 of 3: The only place this analysis exists

Score 69 of 100, ranked 114 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Jason Zhou of Superdesign posted a 40-second film on X calling treg the "OpenSource Clay killer": 85% cheaper than Clay, most accurate on people-search bench, 25x faster, and you run it from Claude Code or Codex. The repository came in the next post. I cloned it, read the routing code and the catalog, pulled the frames out of the film, and checked each number against whatever primary source exists for it.

superdesigndev/treg@f93ae9b · snapshot 2026-10-06
tracked files
4,090
license
Apache-2.0
branch
main
tests
249 files
source
11.8 MB
commit date
2026-10-06
source by language
Python8.6 MB(508)HTML1.5 MB(29)JavaScript666.5 kB(80)Vue474.4 kB(68)CSS387.4 kB(12)TypeScript144.5 kB(43)Shell32.0 kB(10)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

Read at f93ae9b (2026-10-06). Python package tools-registry 0.22.0; 4,090 tracked files, about 90,000 lines of Python under src/treg.

local clone, 2026-10-06 at f93ae9b — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

The short version: the thing that does the work is a small, well-reasoned router, about 1,000 lines across plan.py and route.py, sitting on a large hand-curated price catalog. The router deserves the attention. The launch numbers mostly cannot be checked, and the one chart that can be checked is labelled wrong.

treg's README banner: a dark treg badge, the headline 'OpenRouter for agent tools', the line 'One unified key for 2000+ tools, priced per call, open source', a masked trg_live_ token, and logo pills for Semrush, Moz, Majestic, SerpApi, DataForSEO, SE Ranking, TikTok, Instagram, YouTube, X, Reddit, LinkedIn, Hunter, Lusha, PDL, Apollo, Crunchbase, Coresignal, Google Ads, Meta Ads, GA4, Pinterest and Snapchat, ending in '+2,000 endpoints'.
How treg describes itself: one token in front of thousands of paid endpoints (the project's README, docs/assets/treg-hero.png).
Repositorysuperdesigndev/treg, Python 3.12+, FastAPI server and a treg CLI
Commit readf93ae9b, 2026-10-06
Catalog3,851 endpoints in 109 provider files under src/treg/catalog/ (measured)
Routing90 capability contracts, 449 adapters (measured)
LicenceApache 2.0 plus "Additional Terms": no hosted, managed or embedded service for third parties without written permission
Hostedtreg.to, prepaid balance, $1.00 free for a new team

What treg is

treg is not a people database. It holds no contacts. It is three things stacked on each other.

A catalog. One YAML file per provider (Hunter, Tomba, Apollo, Lusha, Exa, DataForSEO, TikHub and 100 more), each endpoint carrying its parameters, an example response captured live, and a cost block. The cost block is the interesting part: a type (per_success, per_call, per_result, free), a value in the provider's own unit, and a note on how that was established. Most of them read like lab notes. Tomba's says a repeat of the same query within a month bills zero, "confirmed live" six minutes apart; Lusha's says the docs claim search is free and the billing field says one credit per 25 rows. Credits become dollars through fx.yaml, which divides each provider's cheapest public plan by its credits (Tomba: $89 for 10,000 credits, $0.0089 a credit). That makes these list-price upper bounds, as the file itself says.

A proxy that injects credentials. You build the real upstream request and prefix it with https://treg.to/call/. treg resolves the tool by host, adds the key server-side, and relays everything else byte for byte, stripping only hop-by-hop headers, its own X-Treg-* headers and its session cookie. Whose key gets used follows a fixed ladder (README, "the credential ladder"):

  1. a tool your team registered for that provider;
  2. a secret your team stored for it;
  3. a verified public route that needs no key, which is free;
  4. otherwise treg's own key, billed to the team's prepaid balance.

Your own key always wins, and calls on it are never metered.

A meter. A platform-key call reserves its estimated cost from the balance and settles afterwards from what the provider actually reported. Statuses that mean "our key was rejected or throttled" (401, 402, 403, 429 and friends) are never charged. An endpoint with no published price is refused, not served free. The margin over the provider's rate is a deploy setting, platform_margin, whose default in config.py is 0.0; that matches the site's "0% markup", but what the hosted deployment sets is not in the repo.

That is the general product: "OpenRouter for tools". People search is one aisle of it, and the only aisle the launch talks about.

The top of treg.to/people-search: the headline 'Hermes for people search', a line saying one skill gives an agent 1B+ contacts across Apollo, Hunter, Tomba and People Data Labs with 109 providers behind one token, a link reading '27-person receipt: 20 verified emails for $0.58', and a demo table of companies (Modal, Baseten, Fireworks, Together, Modular) with a side panel labelled 'treg · people.email.find, cheapest first, misses free' showing Tomba.
The landing page the post links to, rendered by me in headless Chromium on 2026-10-06. The 109-provider count it prints matches the catalog (treg.to/people-search).

How a routed people search works

Calling one provider is easy. The hard part of "find this person's work email" is that 23 providers can answer it, each wants a different request shape, each bills differently, and each misses on different people. Clay's waterfall column is a hand-ordered list of them. treg's answer is a routed endpoint, POST /call/treg.people.email.find, built from two data files.

A contract (catalog/contracts.yaml) says what the question is, in provider-free terms. For people.email.find it lists four identity variants: {full_name, domain}, {first_name, last_name, domain}, {linkedin_url} and {linkedin_handle}, plus derive rules so a caller may send whichever it holds (first_name is split_first(full_name), a handle is pulled from a URL). It names the output core every answer must fill (email, required; confidence, verified) and what a miss looks like: email == null.

An adapter per endpoint (catalog/adapters.yaml) maps the contract onto that provider: which variants it accepts, where each field goes (full_name → body.data.full_name for Prospeo), where the answer comes back (person.email.email), and its own miss predicate. Adapters are verified at catalog load by round-tripping the endpoint's captured test request and example response; one that fails is never a routing candidate. At this commit 23 adapters route people.email.find and 23 route people.search (measured).

The router then does four things.

1. Keep only candidates that can answer. An adapter that does not accept the identity you sent is dropped with a reason ("needs {linkedin_url}"). For people.search there is a sharper rule: company_domain is a scoping key, so a provider that cannot send it is dropped outright. The comment explains why: a title-only search asked for the CEO of one company "returns CEOs of any company", and each of those bills as a hit.

2. Rank them. rank() in routing/plan.py sorts by a tuple:

# src/treg/domain/catalog/routing/plan.py, rank() — the sort key
(tier,                      # own tool/credential 0, anonymous route 1, treg's key 2
 pref,                      # X-Treg-Route-Prefer, or the contract's default order
 -covers(variant),          # how many of the caller's keys this variant uses
 len(c.ignored),            # filters the caller sent that this adapter drops
 c.expected_cost_per_hit,   # price x P(billed) / P(hit)
 p50_ms, last_ok_days, endpoint_id)

The cost term is the one worth stopping on:

cost per hit=price×P(billed)P(hit)\text{cost per hit} = \frac{\text{price} \times P(\text{billed})}{P(\text{hit})}

Here P(hit)P(\text{hit}) is the endpoint's measured hit rate once it has at least 50 samples, else its success rate, else 1. P(billed)P(\text{billed}) equals P(hit)P(\text{hit}) when the endpoint is priced per_success and 1 otherwise. So for a provider that charges only when it finds something, the hit rate cancels and the expected cost per hit is just its price. For one that charges every call, a 50% hit rate doubles it. Every one of the 13 name-and-domain email finders in the catalog is per_success (measured), which means that for the commonest request the ranking is pure price, and live hit rates only break ties through latency. That is a reasonable outcome, and a consequence the code never states.

The len(c.ignored) term sits ahead of price on purpose. Its comment records the incident: a search for {q, title, location: London, country: GB} went to the cheapest child, which mapped neither geo filter, and "returned people in Bengaluru and San Francisco — reported as a hit, $0.0025". A provider that would answer a looser question now ranks last among equals, and the response names what it ignored in X-Treg-Ignored-Filters, or refuses with an unbilled 422 if you send X-Treg-Route-Strict-Filters: 1.

3. Walk the ranking as a waterfall. route.py calls each candidate through the same execute_call path as any proxied request, so holds, audit rows and settlement are the ordinary ones. On an error (a 5xx, 429 or 402 from the vendor) it tries the next, at most two extra. On a miss it continues, unless you sent X-Treg-Route-Waterfall: 0. On a hit it stops. A candidate whose price would push the running spend past X-Treg-Route-Max-Cost is skipped; the default is $1.00, "a runaway guard, not a budget". For list answers, X-Treg-Route-Min-Results: N marks a short answer as weak and keeps looking, bounded at two extra providers, because unbounded it took the benchmark's lookup set from $1.76 to $22.35 over 28 queries "for answers that were already right".

4. Say what happened. The answer carries the contract's core output, the provider's raw body, and _treg: {served_by, tried} with each attempt's outcome and charge. A hit whose verified is not true also carries _treg.advice telling the agent to run treg.people.email.verify before sending. The contract comment says why: one team "bounced on 73 of 79 addresses" that were unverified directory rows.

The widget replays this with the catalog's own prices. Pick the identity, slide the rung that actually holds the person's email, and toggle an own key.

13 candidates · max cost $1.000

Every name-and-domain finder here is priced per success, so its hit rate cancels out of the ranking. Switch to the LinkedIn identity to see the one that is not.

  1. 1quickenrich.people.email.find$0.0048 · miss
  2. 2trykitt.people.email.find$0.0050 · miss
  3. 3tomba.people.email.find$0.0089 · miss
  4. 4moltsets.people.email.find.name$0.010 · hit · billed $0.010
  5. 5dropleads.people.email.find$0.018 · not asked
  6. 6findymail.search.name$0.020 · not asked
  7. 7limadata.people.email.find.name$0.020 · not asked
  8. 8datagma.people.email.find$0.023 · not asked
  9. 9hunter.people.email.find$0.025 · not asked
  10. 10leadsforge.people.email.find$0.025 · not asked
  11. 11prospeo.people.email.find$0.025 · not asked
  12. 12leadmagic.people.email.find$0.025 · not asked
  13. 13wiza.people.email.find$0.075 · not asked
Found after 4 calls; the team pays $0.010. Prices are the catalog's list rates at commit f93ae9b (credits converted with its fx.yaml); live hit rates and latencies, which break ties on treg.to, are not in the repo and are left out. Which providers the hosted service holds keys for is also not in the repo.

Three things show up quickly. With a name and a domain, the first five rungs cost under two cents each and a miss on any of them is free, so finding a person on rung four costs that rung's price and nothing else. Connecting your own Hunter key puts Hunter first at zero, whatever it costs Hunter's other customers. And on the LinkedIn identity the one per_call endpoint, HarvestAPI at $0.02, slides down the ranking as you lower its hit rate, and bills you on every miss when the waterfall passes through it.

Using it

curl -fsSL https://treg.to/install.sh | sh      # CLI + skills
treg login                                      # GitHub, --email, or --token for agents
treg catalog search "find a work email"         # by the job, not the vendor
treg catalog get treg.people.email.find         # the plan, prices, what each child accepts
treg balance

A routed call from an agent, without the CLI:

curl -s https://treg.to/call/treg.people.email.find \
  -H "X-Treg-Token: $TREG_TOKEN" \
  -H "X-Treg-Route-Max-Cost: 0.10" \
  -H "content-type: application/json" \
  -d '{"full_name": "Alexis Ohanian", "domain": "reddit.com"}'

In Claude Code it is /plugin marketplace add superdesigndev/treg then /plugin install treg@treg; the skill walks the agent through the CLI and treg mcp install. The skill file (skills/treg/SKILL.md, 420 lines) is mostly guardrails: do not retry a find, every hit bills; verify every address; a routed search tells you which filters it dropped. The "run GTM in Claude Code or Codex" part of the launch holds: it is a plugin, a skill and an MCP server.

Self-hosting is uv sync, uv run python -m treg upgrade, uv run python -m treg, on SQLite by default. A self-hosted registry has no platform keys, so the catalog becomes a set of adapters for your own keys. That is useful, and it is the only kind of hosting the licence allows without asking.

Checking the launch claims

"85% cheaper than Clay"

The number comes from the film, not from the repository or the landing page.

A frame from the launch film titled '85% cheaper than Clay', subtitled 'cost of 100 correct work emails, at each tool's list price'. Bars: treg $0.58 (highlighted, WINNER), Clay $3.77, Apollo $2.91, Deepline $3.40, Hunter $2.63. Footer: 'same 100 leads, name + domain in, graded against the published address · Oct 2026 · full table at treg.to/people-search'.
The cost claim, at the frame where every counter has settled. These are the publisher's numbers; no run log or table for them is published (launch film in Jason Zhou's X post, 0:09).

The arithmetic inside the frame holds: 1 − 0.58 / 3.77 = 84.6%, which rounds to 85% (reasoned). The film's closing table gives the same five tools as a cost per 100 leads: treg $0.53 at 92 correct, Clay "~$3.35" at 84 correct. Dividing cost by correct answers reproduces the chart for treg (0.53 / 0.92 = $0.58), Apollo and Hunter, but gives $3.99 for Clay rather than $3.77 (reasoned). The tilde is the film's own: Clay's figure is an estimate "at each tool's list price", not a bill. On either Clay number the saving is about 85%.

What I cannot check is anything underneath. The footer says "full table at treg.to/people-search". On 2026-10-06 that page has no such table, no list of the 100 leads, and no mention of Clay's result beyond its $149/mo plan price (measured, by fetching and grepping it). The page's own self-playing demo meters its 100-lead run to $1.76, a third number for what reads as the same film.

What is checkable is the landing page's smaller receipt: "27 people, 20 verified emails, $0.58". Its CSV is in the repo (src/treg/workflow_runs/find-and-verify-a-lead-list.csv), and it reproduces. 50 companies, 27 kept by the fit gate, 18 emails from Hunter and 3 from Kitt on Hunter's misses, 6 found nowhere, 20 verified valid and 1 unknown (measured). At catalog prices that is 18 × $0.0245 + 3 × $0.005 = $0.456 to find (reasoned). The page adds $0.125 to verify, which is exactly 20 definitive checks at LeadMagic's $0.00625 with the unknown verdict free; the CSV does not name the verifier, so that match is mine. Total $0.58, or $0.029 per usable email (reasoned). It is a different run from the film's, and the coincidence of the two $0.58 figures is worth knowing before quoting either.

One smaller staleness: the page calls $0.0089 (Tomba) the "cheapest email-find on the catalog today". It is hard-coded in the HTML. The catalog at this commit has Kitt at $0.005 per found email and QuickEnrich at $0.004834 (measured).

"25x faster"

A frame from the launch film titled '25x faster than Clay', subtitled 'median time to a found work email'. On the left a faded Clay panel with a timer at 10.9 s and five names still 'finding…'; on the right a highlighted treg panel at 0.44 s listing five names with emails such as maya@baseten.co and sam@dagster.io. Footer: 'same 292 leads, name + domain in · full table at treg.to/people-search'.
The speed claim. Clay's timer is animated and stops at 11.2 s a few frames later (launch film in Jason Zhou's X post, 0:15).

Clay's counter stops at 11.2 s; treg's reads 0.44 s. 11.2 / 0.44 = 25.5, so "25x" is the film's own ratio (reasoned). It is a median over 292 leads, a different set from the cost chart's 100, and nothing else about the run is published. The direction is believable from the mechanism: one synchronous API call to a cheap finder against a spreadsheet tool that queues each row. The factor is the publisher's.

"Most accurate on people-search bench"

There are two accuracy claims, and they are not the same test.

The film's is its own: 92 of 100 leads with a known answer, against Clay 84, Hunter 84, Deepline 80 and Apollo 77 (reported). Like the cost chart, the leads and grading are not published.

A frame from the launch film titled 'Most accurate', subtitled 'correct work emails, 100 leads with a known answer'. Bars: treg 92% (WINNER), Clay 84%, Apollo 77%, Deepline 80%, Hunter 84%.
The film's accuracy chart: treg's own 100-lead test, not the LessieAI benchmark the landing page cites (launch film in Jason Zhou's X post, 0:12).

The landing page's is the named benchmark: "People Search Bench by LessieAI — 119 real tasks, % answered correctly. Same agent, with and without treg", with a headline duel of 43% for Claude Code alone against 78.2% with treg on B2B prospecting, and four category charts. PeopleSearchBench is real and its repository publishes per-category scores for four platforms.

LessieAI's 'Performance by scenario' chart: horizontal bars per scenario. Influencer/KOL: Lessie 62.3, Claude Code 43.2, Exa 41.6, Juicebox 31.1. Expert/Deterministic: Lessie 70.4, Exa 61.2, Claude Code 57, Juicebox 44.2. B2B Prospecting: Lessie 60.6, Exa 55.2, Juicebox 51.4, Claude Code 43. Recruiting: Lessie 68.2, Juicebox 65.7, Exa 64.7, Claude Code 50.5.
The benchmark's own chart. treg does not appear in it (LessieAI/people-search-bench README, assets/performance-by-scenario.png).

I recomputed treg's competitor bars from LessieAI's summary.json files. Every one reproduces as the plain mean of three dimensions: Lessie's B2B 60.6 is (62.8 + 63.5 + 55.5) / 3, Claude Code's 43.0 is (43.0 + 42.3 + 43.6) / 3, and so on for all twelve (measured). That settles what the chart is, and it is not "% answered correctly". The three dimensions are padded nDCG@10 over web-verified relevance grades, an effective-coverage score, and an information-utility score. A 43 does not mean Claude Code answered 43% of B2B tasks correctly.

treg
total only: no breakdown published
78.2
Lessie
60.6
Exa
55.2
Claude Code
43.0
nDCG@10 ÷ 3coverage ÷ 3utility ÷ 3
Each competitor bar is the mean of three 0-100 scores from LessieAI's published summary.json: padded nDCG@10, effective coverage and information utility. None of them is a percentage of tasks answered correctly, which is how treg's page labels the chart. treg's own bar appears in neither repository with its components or a run log.

Three more things the comparison leaves out:

The repository is candid about part of this. The Enrich Arena's benchmark page calls its charts "a published snapshot, not a new benchmark run", and its design doc says "the original repository's current figures differ from the landing" and that the page "does not manufacture an overall score". The landing page's subtitle did not get the same care.

"Open source"

The licence is Apache 2.0 with Additional Terms that "take precedence over the Apache License to the extent of any conflict". The first term forbids using the software "to provide a hosted, managed, or embedded service to third parties" without written authorisation. That is a field-of-use restriction, which the Open Source Definition does not allow, so this is source-available in the usual sense of the words. For a team that wants to run its own registry, nothing is lost: self-hosting for your own organisation is "expressly permitted and encouraged", and calling the hosted API from your own product is allowed.

What the launch leaves out

Superdesign's own marketing draft, marketing/rebuild/01-rebuild-clay.md, is a better review of the Clay comparison than the film. It lists where Clay is still the right buy: you want a spreadsheet, not a script; you need a research column that reads each company's website; you run fifty tables unattended and want Clay's queueing and rate-limit handling: "Here, that's your job." It also says plainly that Clay's moat is "150-plus pre-negotiated data providers". treg's answer to that moat is its catalog: real, long, and priced at list rates it documents per provider.

A reply under the launch post made the remaining point: a people-search score is the easy part, and the test is how stale the data is after 90 days and how many sends bounce. treg holds no data of its own, so staleness is whatever the answering provider's is. Its defence is the verify step it keeps pushing, a fraction of a cent per address, which caught the bounces in its own incident notes.

Verdict

ClaimStatus
Run it from Claude Code or CodexHolds: plugin, skill and MCP server in the repo
85% cheaper than ClayArithmetic consistent within the film; underlying run not published
25x faster11.2 s / 0.44 s in the film; underlying run not published
Most accurate on People Search BenchCompetitor bars reproduce; the label is wrong, Juicebox is omitted, and treg's own scores have no published run
27 people, 20 verified emails, $0.58Reproduces from the committed CSV and catalog prices
Open sourceApache 2.0 with a hosted-service restriction: source-available

The piece I would take from treg is not a number from the film. It is the router: contracts that state the question without naming a vendor, adapters that are verified against captured fixtures before they may serve, a ranking that puts "answers the question you asked" ahead of "is cheapest", and a cost term that is honest about what a miss costs. If you build an agent that calls paid APIs, plan.py is 183 lines and worth reading whole. For a sibling story about benchmark multipliers that come apart into simpler factors, see the WindTunnel decomposition; for another way to keep an agent's paid side effects on a leash, Cloudflare OS's gatekeepers; and for harnesses read for where they check the agent's work, three agent harnesses.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "treg: a people-search router priced by expected cost per hit", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026treg,
  author = {Satyajit Ghana},
  title  = {treg: a people-search router priced by expected cost per hit},
  url    = {https://ai.thesatyajit.com/articles/treg},
  year   = {2026}
}
share