2026-10-06 · 21 min · search · developer-tools · mcp · agents
Why read this
Solidtop 85%Reads OpenSEO's AI-visibility code end to end: who answers, what counts as a mention, and how much noise one answer per prompt leaves in the trend.
- Original analysis
- Concrete numbers to act on
- Open code or weights
Developer tools & infraAPI onlyMITPractitioner tool
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 58 of 100, ranked 272 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
This site spends a fair amount of effort on being readable by machines. Every
article emits a JSON-LD citation list pulled from its body, there is an
/llms.txt and a full-corpus /llms-full.txt, every page has a .md twin, and
robots.txt waves in GPTBot, ClaudeBot, PerplexityBot and ten others by
name. What I have never had is any way to see whether any of it changes what an
answer engine says. So when OpenSEO announced "AI Visibility" on its
$10/month plan, with the line "Instead of paying $100+/month", I cloned the
repository to find out what exactly it measures.
- license
- MIT
- branch
- main
- tests
- 283 files
- source
- 6.7 MB
- commit date
- 2026-10-06
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
Read at deb4491 (2026-10-06). open-seo 0.1.11, MIT; 1,870 tracked files, 222 of them tests. The AI-visibility server code is about 3,300 lines.
local clone, 2026-10-07 at deb4491 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history

What surprised me is how little of it OpenSEO does itself, and how sensible that is. The answers are collected by DataForSEO, a data vendor that runs the consumer ChatGPT and Gemini apps and Google's results page and sells you the output at a tenth of a cent each. OpenSEO's contribution is the part after that: deciding whether an answer names you, whether it links to you, and whether the number on the chart moved for a reason. That part is small, about 250 lines for parsing and matching, and it is careful in ways I did not expect. The weakness is not in the code. It is that each check asks every question exactly once.
Three pages, three different sources
The launch posts describe one feature. The code has three, and they do not share a data source, so the answers on each page mean different things.
Prompt Research starts from a keyword ("password manager") and lists
questions people ask about it. I assumed this would be some sample of real
ChatGPT prompts. It is not, and to OpenSEO's credit their own docs say so: the
questions come from DataForSEO's dataset, "built mostly from Google 'People also
ask' questions; they are not logged ChatGPT prompts". The code calls DataForSEO's
llm_mentions search with a limit of 100 rows, fixed to US English
(CHATGPT_LOCATION_CODE = 2840), ordered by an ai_search_volume that a comment
in aiPromptResearch.ts describes as "estimated from Google 'People also ask'
data, not counted from AI usage". OpenSEO uses it only for ordering.
Next to that it fetches the live Google results for the keyword, 50 deep, keeps the registrable domains that rank (minus a hand list of 18 sites like Reddit, YouTube and Wikipedia that rank for everything), and uses them to decide which questions are on topic: a question survives if it contains the keyword as a phrase of two or more words, or if its answer leaned on a site that ranks for the keyword. Near-duplicates merge after dropping years and articles, so "best password manager in 2025" and "the best password manager 2026" become one row. It is a good filter for a cheap input. You can see where it leaks in the launch screenshot: "What's the safest email account to have?" and "Is there a better email than gmail?" are in a password-manager list, presumably because their answers cite a privacy-and-security site that also ranks for the keyword.

Prompt Explorer asks one prompt to ChatGPT, Claude, Gemini and Perplexity
through DataForSEO's llm_responses endpoints, which are API calls rather than
the consumer apps. The model names are not hardcoded. llm-models.ts reads
DataForSEO's free model catalog each hour and picks the newest plain alias per
family, with ChatGPT pinned to gpt-5.6-luna; the fallback list names
claude-sonnet-5, gemini-2.5-pro and sonar-reasoning-pro. The comment
explaining why is worth reading, because it records a real failure: "we shipped
gpt-5 while the catalog had gpt-5.5", and DataForSEO bills a task that fails on
an unknown model name, so every name is checked against the catalog before a
paid call.

Prompt Tracking is the part the launch is about, and the only one that
reads the consumer products. ai-tracking.ts maps each engine to an endpoint:
// src/server/lib/dataforseo/ai-tracking.ts:16-20
const AI_TRACKING_ENDPOINTS: Record<AiEngine, string> = {
chatgpt: "/v3/ai_optimization/chat_gpt/llm_scraper",
gemini: "/v3/ai_optimization/gemini/llm_scraper",
google_ai_overview: "/v3/serp/google/organic",
};ChatGPT and Gemini come from DataForSEO's LLM Scraper. Google AI Overviews is
an ordinary organic results page, ten deep, with load_async_ai_overview: true
so DataForSEO waits for the overview Google paints after the page loads; the
overview is then fished out of the result items by type. Claude and Perplexity
are not trackable at all. They exist only in the Explorer, through the API.
What I could not check is what "the consumer version" means on DataForSEO's side: logged out or in, which model tier, whether web search is on, what location the browser claims. DataForSEO's pricing page does not say, and nothing in OpenSEO controls it beyond the market (country and language) passed with each task. Every number on the tracking page inherits those unknowns.
One check, from post to chart
A check is a Cloudflare Workflow, AiVisibilityWorkflow.ts, and reading it
tells you where the money goes. It does four things.
It prepares: marks the run as running, and on the hosted service refuses to
start unless the organisation's credits cover the whole check. It posts:
groups the pending answers by engine into batches of up to 100 and sends each
batch as one DataForSEO task_post. That call is the one that costs money, so
it runs with zero retries and maxServerErrorRetries: 0, with the comment "A
billed task_post must never be replayed on an ambiguous 5xx." Then it polls.
Reading a finished task is free, so the collect step may retry, and the
workflow sleeps on a schedule whose cumulative waits are 2, 4, 7, 10, 15, 20,
30, 45 and 60 minutes. DataForSEO's standard queue promises results within 45
minutes, so the last poll leaves a quarter of an hour of slack. Anything still
missing after that is marked failed, and the run ends as completed, partial
or failed.
Each answer that arrives goes through parseDataforseoAnswer, which keeps the
answer markdown and the cited sources, de-duplicates the URLs, strips utm_*,
gclid, fbclid and msclkid, and numbers them in the engine's order. No
page is fetched and no model is called at this stage; the comment says all
evidence comes from the captured result. Then every tracked brand, your own
plus each competitor on the project, is matched against it.
What counts as a mention
The matching code is where I expected shortcuts and found mostly care. There are two tests, and they are deliberately separate.
Cited is the simple one: any source URL whose host equals your domain or ends
in .yourdomain, ignoring www.. A link to help.bitwarden.com counts for
bitwarden.com; a link to bitwarden.com.evil.example does not.
Mentioned is a regex over the answer's prose, but only after two cleanups. The
first is markdownToText, which deletes citation pills. If the answer contains
[bitwarden.com](https://bitwarden.com/help/) or [1](…), that link label is
evidence of a source, not a textual mention, so it is removed before matching.
The second blanks out every URL and email address in the prose, keeping their
length so the character offsets still point into the original answer. Then the
brand's name and its domain are searched case-insensitively with Unicode word
boundaries, with any run of whitespace in the name allowed to match any run in
the answer, and both sides normalised to NFD so Café matches either encoding
of the accent. Overlapping matches keep the longest.
The result has three fields: mentioned, cited, and firstMention, the
character offset of the first match. The widget below runs the same functions,
ported line for line, on four answers I wrote to show the edges.
For a free option, 1Password has no free tier, so most reviewers point to an open-source manager with unlimited devices [1] and a self-hostable server . Proton Pass is the other common pick.
Three things fall out of it, and all three are visible in the code rather than hidden.
The brand name is the project name, and there is exactly one. aiBrands() in
aiVisibilityConfiguration.ts builds the own brand from project.name and
project.domain, and competitors from the project's competitor list, each with
one name. There is no alias list. If people call you by a short name, a
former name or a product name, those answers do not count unless they also
spell out the domain.
A generic name matches its dictionary word. The function's docstring says so plainly: "A generic brand name can count as a mention of something else; that is accepted for simplicity." The release notes for 0.1.11 repeat it as a beta warning. If your company is called Linear, Notion or Apple, the mention rate is an upper bound.
And "average position" on the Competitors tab is not a rank in the answer. In
summarizeAiBrands, each answer's brands are sorted by their first-mention
offset and numbered, but only the brands you are tracking are in that list. A
position of 1.7 means "usually named first or second among the six brands I
told it about", not "second recommendation in the answer". An answer that
opens with five products you did not list still puts your brand at 1.
The numbers on the launch dashboard
The launch post follows Bitwarden through the product, and the Competitors screenshot is the most quoted image from it: Bitwarden named in 93% of non-branded answers, 1Password in 78%.

Look at the Prompts tab first, because it shows the sampling plainly. Every
prompt row reads "3 of 3" or "1 of 3". With three engines and one answer per
engine, a prompt's rate can only take four values. The cost estimator in
aiVisibilityCost.ts states the design outright in a warning it shows before
every check: "One fresh answer per prompt and engine."

So the 93% is 25 answers out of 27, and the 78% is 21 out of 27. I put Wilson 95% intervals on both: Bitwarden's runs from 76.6% to 97.9%, 1Password's from 59.2% to 89.4%. They overlap. Bitwarden's citation rate, 13 of 26, is anywhere from 32.1% to 67.9%. None of this means the screenshot is wrong; it is one check, and the blog itself says "judge your visibility by the trend over a few weeks, not by one check". It does mean the gap between first and second place in that table is not something one check can establish.
One small thing I could not reconcile. In the screenshot, mentions are out of
27 and citations out of 26, and on the Prompts tab one topic shows "13 of 15"
beside "6 of 14". At the commit I read, both rates in the Competitors tab
divide by the same summary.answers, and the Prompts tab counts the same
eligible answers for both columns. The screenshots were presumably taken on an
earlier build that counted citations differently; I do not know what it
excluded.
The Citations tab is, to me, the most useful view in the product, because it is a list rather than a rate. Grouped by domain, it shows which third-party pages the engines lean on for a category.

The per-engine split is the interesting part. In this one check ChatGPT cites bitwarden.com 8 times and 1password.com 5 times and never cites allaboutcookies.org, while Google's overviews cite YouTube 9 times. Gemini cites passwordmanager.com 6 times and the overviews never do. Whatever each engine retrieves from, it is not the same index, and "get cited by AI" splits into three separate jobs.
How much one answer per prompt can tell you
The trend chart is where the design earns its keep, and where the sampling
bites. aiVisibilityTrend.ts is the best-reasoned file in the feature. It
splits time into the last 7, 28 or 90 days and the same span before it, and it
does not simply pool the answers. It builds cells, one per prompt, engine,
market and own-brand identity, computes each cell's rate, and averages cells.
Only cells with answers in both periods enter the comparison.
This is a paired design, and the right one. Adding ten new prompts this
week cannot move the comparison, because the new cells have no previous period.
Editing a prompt's wording starts a new prompt, so it does not either. Renaming
your brand starts new cells. Prompts that contain your brand name or domain are
excluded, so "Is Bitwarden safe to use?" does not prop up the rate. Failed and
empty collections never count as absent: addAnswer skips them, and the
comment above it says so. And if fewer than 90% of planned answers arrived in
either period (MIN_COVERAGE = 0.9), the rates still show but the change is
hidden.
What the cells cannot fix is that each one, per check, holds a single yes or no. If the true chance that ChatGPT names you for a prompt is 50%, one check tells you heads or tails. Pooling many prompts and many checks is the only way to shrink that, and the arithmetic is unforgiving: the standard error of a rate over independent yes/no answers is , the difference of two periods has times that, and the smallest change a 95% test separates from noise is about .
Take the launch post's own example, 20 prompts on ChatGPT and Gemini. Checked weekly with the 7-day window, each period holds one check of 40 answers. At a 50% mention rate the per-period error is 7.9 points and the smallest trustworthy week-over-week change is about 21.9 points. Daily checks put 280 answers in each 7-day period, and the band narrows to about 8.3 points; that costs about $2.40 a month on the hosted plan. Weekly checks read over the 28-day window get you 160 answers a period for $0.32 a month and a band of about 11 points, at the price of seeing a change a month late.
Two assumptions in that arithmetic pull opposite ways, and the widget's caption says which. Treating every prompt as having the same rate is the worst case for a given average; real prompts that are nearly always yes or always no contribute less noise. Treating runs as independent is the best case. If DataForSEO's scraper or the engine returns much the same answer on consecutive days, the effective sample is smaller than the count. I have no way to measure that correlation from outside, so I read the band as an order of magnitude.
The Explorer has a related trap. Rerunning the same prompt with the same
models within seven days is free, which the blog presents as a perk. In
promptExplorer.ts that is PROMPT_RESPONSE_TTL_SECONDS = 7 * 24 * 60 * 60:
the second run is the first run's answer from the cache. It is the right
billing choice, and it means the Explorer cannot show you variance at all.
If you want to see how much an answer moves, rephrase the prompt by one word.
Nothing here is a defect in OpenSEO. One answer per prompt and engine is what $0.002 buys, and the tool tells you so before every check. The trouble is the dashboard: a column of percentages with no interval invites you to read a 5-point move as news. I would like to see the interval drawn on the trend chart; the code already has every number needed to draw it.
What $0.002 per answer is made of
The blog's price is "$0.002 per answer". The constant behind it is
AI_RECORD_COST_USD = 0.0012 in src/shared/ai-visibility.ts, which matches
DataForSEO's published standard-queue price of $0.0012 per LLM Scraper result.
For AI Overviews the comment splits it into a results page at $0.0006 plus the
async overview load at $0.0006, and DataForSEO refunds the second half when
Google shows no overview.
On the hosted service, src/shared/billing.ts applies SEO_DATA_COST_MARKUP = 1.28, converts to credits at 1,000 per dollar, and rounds up. One answer is
0.0012 × 1.28 = $0.001536, which ceils to 2 credits, or $0.002. This per-answer figure is the estimate, and the number the blog quotes. The settlement is different.
settleUsageCredits charges each provider call separately on DataForSEO's
reported cost, and one task_post carries a whole batch of up to 100 answers.
Twenty ChatGPT answers settle at 31 credits, not 40. At batch sizes like that
the hosted price lands at about $0.0015 per answer, which is the README's
"28% extra for every request" and not the 67% that the per-answer estimate
implies. So the estimate is honest in the direction that matters: it
overstates, and the credit check before a run uses it, so a run never starts on
credits that cannot cover it.
The blog's examples check out against those rules. Twenty prompts on two
engines is 40 answers, $0.08 per check estimated; weekly is about $0.32 a
month and daily about $2.40, both at the planning figures of 4 and 30 checks a
month in scheduledChecksPerMonth. Self-hosted, the same 40 answers cost
$0.048 paid straight to DataForSEO. Prompt Research is the expensive call: a
llm_mentions request at $0.1 plus $0.001 per row, up to 100 rows, plus a
50-deep live results page at $0.008. At the full 100 rows that is $0.208
raw and about $0.27 with the markup, close to the docs' "about $0.25 per
keyword". It is cached for a day per organisation.
I did not check the "$100+/month" for incumbent AI-visibility products. It is plausible for the enterprise tools, but I have no primary source in front of me, so I leave it as the launch's claim.
Running it yourself
The licence is MIT, copyright Ben Senescu, with no extra terms. It is cleaner than most "open-source alternative" launches I have read; the treg teardown, for contrast, found an Apache 2.0 licence with a clause forbidding hosting.
There are two self-hosting paths. Docker runs the Cloudflare Workers runtime
(workerd) inside a Node 22 image, keeps state in a volume, and runs with
AUTH_MODE=local_noauth, which means no login at all; the docs say to put it
behind your own authenticated proxy. A built-in scheduler looks for due checks
every five minutes, so scheduled tracking works with no host cron. The
recommended path is Cloudflare itself, which the README says works on the free
plan. Either way you need a DataForSEO API key, which is a paid account, so
runsOn for this page is honestly "an API", not "your laptop". Setup research
and generated prompt suggestions also want an OPENROUTER_API_KEY; the
default model is openai/gpt-5.6-luna, overridable with OPENROUTER_MODEL.
Two defaults are worth knowing before you run it. Telemetry is on: heartbeats
with aggregate counts tied to a random install ID, every five minutes for the
first two hours and then at most daily, documented as never including URLs,
keywords or prompts. OPENSEO_TELEMETRY_DISABLED=1 turns it off. And tracking
cannot collect ChatGPT answers for one country in the shared market list,
location 2275 (Palestine), per AI_UNSUPPORTED_LOCATIONS.
The MCP server deserves a sentence. An agent can research, explore, configure
and read results, but a paid check needs a maxCostUsd the user approved from
a quote, and aiVisibilityRuns.ts refuses the run if the price has since
gone up. Letting an agent spend money is the risky part of agent tooling, and
this is a clean way to bound it. The jev-linkmap
piece is a reminder of how quickly an agent's tuning bill outgrows the
headline price when nothing bounds it.
What this can and cannot see about a site like this one
Back to why I looked. Everything this site does for discoverability is on the
input side. citationsFromBody() in lib/jsonld.tsx puts up to 20 arXiv,
DOI, GitHub and Hugging Face links into each article's citation field.
/llms.txt is an index for agents, the .md twins are clean text, and
robots.txt declares ai-train=yes, search=yes, ai-input=yes for every
crawler. None of that can be observed from the outside except through what
the engines eventually say.
OpenSEO measures exactly that output, for three surfaces, for the questions you choose. For a site like this, the owned-citation rate is the only number that matters: nobody asks ChatGPT about "Satyajit Ghana" as a brand, but someone might ask how SGLang's radix cache works, and the question is whether the answer links here. The Citations tab would show which pages get cited instead, which is a concrete thing to learn.
What it cannot do is connect an input to an output. It never sees whether
llms.txt was fetched, whether a page was in a training set, or which
retrieval index a scraper's session hit. It does not track Claude or
Perplexity at all, and it reads answers from whatever session state the
vendor's scraper has, not a logged-in user with memory and history. The
questions it suggests are People-also-ask questions, so they describe Google
searchers rather than chat users. And the noise band above is the ceiling on
attribution: if I ship a JSON-LD change and the owned-citation rate moves by 6
points on a weekly check, the tool has told me nothing. With daily checks over
a 28-day window, a move of that size starts to mean something, and even then
it says that the answers changed, not why.
Still, it is more than I have now. The thing I would actually do is pick a dozen questions this site's articles answer better than most, track them on ChatGPT and AI Overviews daily (the two engines that, in the launch screenshot, cite the most different sources), and read the Citations tab monthly for which pages win instead. The rate on the chart I would treat as a smoke alarm, not a thermometer.
How I checked
I shallow-cloned every-app/open-seo at deb4491 (2026-10-06) and read the
AI-visibility path end to end: the workflow, the DataForSEO clients and price
table, the parser, the matcher, the trend and results services, the billing
helpers and the relevant UI components. The mention widget is a line-for-line
port of markdownToText, aiMentionSpans and ownedAiDomain, minus the NFD
offset mapping, and I ran the ported functions on its four presets in Node to
confirm the outcomes it shows. I did not run OpenSEO or call DataForSEO. The
Wilson intervals and the noise arithmetic are mine, computed from the counts
in the launch screenshots and from the formula stated above. DataForSEO's
$0.0012 standard-queue price is from its LLM Scraper pricing page. The launch
posts and their thread were read through the fxtwitter mirror; the figures are
the launch card from X and the five screenshots committed with OpenSEO's own
launch post under web/public/blog/ai-visibility-in-openseo/.