# OpenSEO's AI Visibility: a scraper, a regex and one answer per prompt

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/open-seo
> date: 2026-10-06
> tags: search, developer-tools, mcp, agents

This site spends a fair amount of effort on being readable by machines. Every
article emits a JSON-LD `citation` list pulled from its body, there is an
`/llms.txt` and a full-corpus `/llms-full.txt`, every page has a `.md` twin, and
`robots.txt` waves in `GPTBot`, `ClaudeBot`, `PerplexityBot` and ten others by
name. What I have never had is any way to see whether any of it changes what an
answer engine says. So when OpenSEO announced "AI Visibility" on its
\$10/month plan, with the line "Instead of paying \$100+/month", I cloned the
repository to find out what exactly it measures.

<RepoCard repo="every-app/open-seo" note="Read at deb4491 (2026-10-06). open-seo 0.1.11, MIT; 1,870 tracked files, 222 of them tests. The AI-visibility server code is about 3,300 lines." />

<Figure
  src="https://ai.thesatyajit.com/articles/open-seo/fig1.jpg"
  alt="OpenSEO's launch card over a photo of a stormy sea and rocks: 'OpenSEO, Now including AI visibility' on the left; on the right 'Connect your AI agent' with four agent logos, and 'AI visibility features: Prompt Research, Prompt Explorer, Prompt Tracking'."
  caption="The launch card from the announcement on X: three pages, and an MCP server so an agent can drive them (OpenSEO's launch post, 2026-10-06)."
/>

What surprised me is how little of it OpenSEO does itself, and how sensible
that is. The answers are collected by DataForSEO, a data vendor that runs the
consumer ChatGPT and Gemini apps and Google's results page and sells you the
output at a tenth of a cent each. OpenSEO's contribution is the part after
that: deciding whether an answer names you, whether it links to you, and
whether the number on the chart moved for a reason. That part is small, about
250 lines for parsing and matching, and it is careful in ways I did not expect.
The weakness is not in the code. It is that each check asks every question
exactly once.

## Three pages, three different sources

The launch posts describe one feature. The code has three, and they do not
share a data source, so the answers on each page mean different things.

Prompt Research starts from a keyword ("password manager") and lists
questions people ask about it. I assumed this would be some sample of real
ChatGPT prompts. It is not, and to OpenSEO's credit their own docs say so: the
questions come from DataForSEO's dataset, "built mostly from Google 'People also
ask' questions; they are not logged ChatGPT prompts". The code calls DataForSEO's
`llm_mentions` search with a limit of 100 rows, fixed to US English
(`CHATGPT_LOCATION_CODE = 2840`), ordered by an `ai_search_volume` that a comment
in `aiPromptResearch.ts` describes as "estimated from Google 'People also ask'
data, not counted from AI usage". OpenSEO uses it only for ordering.

Next to that it fetches the live Google results for the keyword, 50 deep, keeps
the registrable domains that rank (minus a hand list of 18 sites like Reddit,
YouTube and Wikipedia that rank for everything), and uses them to decide which
questions are on topic: a question survives if it contains the keyword as a
phrase of two or more words, or if its answer leaned on a site that ranks for
the keyword. Near-duplicates merge after dropping years and articles, so "best
password manager in 2025" and "the best password manager 2026" become one row.
It is a good filter for a cheap input. You can see where it leaks in the
launch screenshot: "What's the safest email account to have?" and "Is there a
better email than gmail?" are in a password-manager list, presumably because
their answers cite a privacy-and-security site that also ranks for the keyword.

<Figure
  src="https://ai.thesatyajit.com/articles/open-seo/fig2.png"
  alt="Prompt Research for the keyword 'password manager' in a Bitwarden project: a list of questions such as 'Which password manager has never been hacked?', 'What is the downside of 1password?', 'Is bitwarden still good in 2026?', 'What's the safest email account to have?' and 'Does nasa use bitwarden?', each with a count of 2 to 6 sources; four carry a green You badge; two are ticked and a 'Track selected (2)' button is active."
  caption="Prompt Research: People-also-ask questions filtered to the keyword, with the sources DataForSEO's stored ChatGPT answer cited. 'You' means that answer named the project's brand (OpenSEO launch blog post, prompt-research.png)."
/>

Prompt Explorer asks one prompt to ChatGPT, Claude, Gemini and Perplexity
through DataForSEO's `llm_responses` endpoints, which are API calls rather than
the consumer apps. The model names are not hardcoded. `llm-models.ts` reads
DataForSEO's free model catalog each hour and picks the newest plain alias per
family, with ChatGPT pinned to `gpt-5.6-luna`; the fallback list names
`claude-sonnet-5`, `gemini-2.5-pro` and `sonar-reasoning-pro`. The comment
explaining why is worth reading, because it records a real failure: "we shipped
gpt-5 while the catalog had gpt-5.5", and DataForSEO bills a task that fails on
an unknown model name, so every name is checked against the catalog before a
paid call.

<Figure
  src="https://ai.thesatyajit.com/articles/open-seo/fig3.png"
  alt="Prompt Explorer with the prompt 'What's the best free password manager?', highlight brand Bitwarden, all four models ticked (ChatGPT, Claude, Gemini, Perplexity) and web search allowed. The first result card is ChatGPT, model gpt-5.6-luna, with a green Bitwarden badge, 369 tokens, and the answer 'Bitwarden is the best free password manager for most people', followed by a bulleted list of reasons."
  caption="Prompt Explorer: one prompt, four API models, the brand flagged where it appears. These are API answers, which can differ from the consumer apps that tracking reads (OpenSEO launch blog post, prompt-explorer.png)."
/>

Prompt Tracking is the part the launch is about, and the only one that
reads the consumer products. `ai-tracking.ts` maps each engine to an endpoint:

```ts
// src/server/lib/dataforseo/ai-tracking.ts:16-20
const AI_TRACKING_ENDPOINTS: Record<AiEngine, string> = {
  chatgpt: "/v3/ai_optimization/chat_gpt/llm_scraper",
  gemini: "/v3/ai_optimization/gemini/llm_scraper",
  google_ai_overview: "/v3/serp/google/organic",
};
```

ChatGPT and Gemini come from DataForSEO's LLM Scraper. Google AI Overviews is
an ordinary organic results page, ten deep, with `load_async_ai_overview: true`
so DataForSEO waits for the overview Google paints after the page loads; the
overview is then fished out of the result items by type. Claude and Perplexity
are not trackable at all. They exist only in the Explorer, through the API.

What I could not check is what "the consumer version" means on DataForSEO's
side: logged out or in, which model tier, whether web search is on, what
location the browser claims. DataForSEO's pricing page does not say, and
nothing in OpenSEO controls it beyond the market (country and language) passed
with each task. Every number on the tracking page inherits those unknowns.

## One check, from post to chart

A check is a Cloudflare Workflow, `AiVisibilityWorkflow.ts`, and reading it
tells you where the money goes. It does four things.

It prepares: marks the run as running, and on the hosted service refuses to
start unless the organisation's credits cover the whole check. It posts:
groups the pending answers by engine into batches of up to 100 and sends each
batch as one DataForSEO `task_post`. That call is the one that costs money, so
it runs with zero retries and `maxServerErrorRetries: 0`, with the comment "A
billed task_post must never be replayed on an ambiguous 5xx." Then it polls.
Reading a finished task is free, so the collect step may retry, and the
workflow sleeps on a schedule whose cumulative waits are 2, 4, 7, 10, 15, 20,
30, 45 and 60 minutes. DataForSEO's standard queue promises results within 45
minutes, so the last poll leaves a quarter of an hour of slack. Anything still
missing after that is marked failed, and the run ends as completed, partial
or failed.

Each answer that arrives goes through `parseDataforseoAnswer`, which keeps the
answer markdown and the cited sources, de-duplicates the URLs, strips `utm_*`,
`gclid`, `fbclid` and `msclkid`, and numbers them in the engine's order. No
page is fetched and no model is called at this stage; the comment says all
evidence comes from the captured result. Then every tracked brand, your own
plus each competitor on the project, is matched against it.

## What counts as a mention

The matching code is where I expected shortcuts and found mostly care. There
are two tests, and they are deliberately separate.

Cited is the simple one: any source URL whose host equals your domain or ends
in `.yourdomain`, ignoring `www.`. A link to `help.bitwarden.com` counts for
`bitwarden.com`; a link to `bitwarden.com.evil.example` does not.

Mentioned is a regex over the answer's prose, but only after two cleanups. The
first is `markdownToText`, which deletes citation pills. If the answer contains
`[bitwarden.com](https://bitwarden.com/help/)` or `[1](…)`, that link label is
evidence of a source, not a textual mention, so it is removed before matching.
The second blanks out every URL and email address in the prose, keeping their
length so the character offsets still point into the original answer. Then the
brand's name and its domain are searched case-insensitively with Unicode word
boundaries, with any run of whitespace in the name allowed to match any run in
the answer, and both sides normalised to NFD so `Café` matches either encoding
of the accent. Overlapping matches keep the longest.

The result has three fields: `mentioned`, `cited`, and `firstMention`, the
character offset of the first match. The widget below runs the same functions,
ported line for line, on four answers I wrote to show the edges.

<MentionMatcher />

Three things fall out of it, and all three are visible in the code rather than
hidden.

The brand name is the project name, and there is exactly one. `aiBrands()` in
`aiVisibilityConfiguration.ts` builds the own brand from `project.name` and
`project.domain`, and competitors from the project's competitor list, each with
one name. There is no alias list. If people call you by a short name, a
former name or a product name, those answers do not count unless they also
spell out the domain.

A generic name matches its dictionary word. The function's docstring says so
plainly: "A generic brand name can count as a mention of something else; that
is accepted for simplicity." The release notes for 0.1.11 repeat it as a beta
warning. If your company is called Linear, Notion or Apple, the mention rate is
an upper bound.

And "average position" on the Competitors tab is not a rank in the answer. In
`summarizeAiBrands`, each answer's brands are sorted by their first-mention
offset and numbered, but only the brands you are tracking are in that list. A
position of 1.7 means "usually named first or second among the six brands I
told it about", not "second recommendation in the answer". An answer that
opens with five products you did not list still puts your brand at 1.

## The numbers on the launch dashboard

The launch post follows Bitwarden through the product, and the Competitors
screenshot is the most quoted image from it: Bitwarden named in 93% of
non-branded answers, 1Password in 78%.

<Figure
  src="https://ai.thesatyajit.com/articles/open-seo/fig4.png"
  alt="The Prompts tab of Prompt Tracking for Bitwarden. Two topics: 'Switching and teams' with brand mentions 100% (15 of 15) and owned citations 67% (10 of 15), and 'Choosing a password manager' with 87% (13 of 15) and 43% (6 of 14). Each prompt row lists ChatGPT, Gemini and Google AI Overviews and shows rates out of 3, such as 100% 3 of 3 mentions and 33% 1 of 3 citations. 'Is Bitwarden safe to use?' is marked Branded."
  caption="Prompt Tracking, one check: each prompt is asked once per engine, so a prompt's rate can only be 0, 33, 67 or 100% (OpenSEO launch blog post, prompt-tracking.png)."
/>

Look at the Prompts tab first, because it shows the sampling plainly. Every
prompt row reads "3 of 3" or "1 of 3". With three engines and one answer per
engine, a prompt's rate can only take four values. The cost estimator in
`aiVisibilityCost.ts` states the design outright in a warning it shows before
every check: "One fresh answer per prompt and engine."

<Figure
  src="https://ai.thesatyajit.com/articles/open-seo/fig5.png"
  alt="The Competitors tab for Bitwarden across 3 engines: Bitwarden (You) mentioned in 25 of 27 answers (93%), cited in 13 of 26 (50%), average position 1.7; 1Password 21 of 27 (78%), 9 of 26 (35%), 1.6; Dashlane 7 of 27 (26%), 1 of 26 (4%), 3.9; Keeper 6 of 27 (22%), 0 of 26, 3.0; NordPass 10 of 27 (37%), 0 of 26, 3.1; Proton Pass 16 of 27 (59%), 5 of 26 (19%), 2.8. Below, suggested competitors: allaboutcookies.org cited in 12 answers, security.org 7, passwordmanager.com 6, pcmag.com 6."
  caption="The Competitors tab from the launch: 27 non-branded answers in one check, consistent with nine prompts on three engines (OpenSEO launch blog post, competitors.png)."
/>

So the 93% is 25 answers out of 27, and the 78% is 21 out of 27. I put Wilson
95% intervals on both: Bitwarden's runs from 76.6% to 97.9%, 1Password's from
59.2% to 89.4%. They overlap. Bitwarden's citation rate, 13 of 26, is anywhere
from 32.1% to 67.9%. None of this means the screenshot is wrong; it is one
check, and the blog itself says "judge your visibility by the trend over a few
weeks, not by one check". It does mean the gap between first and second place
in that table is not something one check can establish.

One small thing I could not reconcile. In the screenshot, mentions are out of
27 and citations out of 26, and on the Prompts tab one topic shows "13 of 15"
beside "6 of 14". At the commit I read, both rates in the Competitors tab
divide by the same `summary.answers`, and the Prompts tab counts the same
eligible answers for both columns. The screenshots were presumably taken on an
earlier build that counted citations differently; I do not know what it
excluded.

The Citations tab is, to me, the most useful view in the product, because
it is a list rather than a rate. Grouped by domain, it shows which third-party
pages the engines lean on for a category.

<Figure
  src="https://ai.thesatyajit.com/articles/open-seo/fig6.png"
  alt="The Citations tab grouped by domain across 3 engines: bitwarden.com, your domain, cited in 16 answers across 9 prompts (ChatGPT 8, Gemini 1, Google AI Overviews 7); allaboutcookies.org, third party, 12 answers (Gemini 4, AI Overviews 8, none from ChatGPT); youtube.com 9 answers, all AI Overviews; reddit.com 8; 1password.com, competitor, 7 (ChatGPT 5); security.org 7; passwordmanager.com 6, all Gemini; pcmag.com 6."
  caption="The Citations tab: which domains the answers cited, and on which engine. The three engines lean on different sources (OpenSEO launch blog post, citations.png)."
/>

The per-engine split is the interesting part. In this one check ChatGPT cites
bitwarden.com 8 times and 1password.com 5 times and never cites
allaboutcookies.org, while Google's overviews cite YouTube 9 times. Gemini cites
passwordmanager.com 6 times and the overviews never do. Whatever each engine
retrieves from, it is not the same index, and "get cited by AI" splits into
three separate jobs.

## How much one answer per prompt can tell you

The trend chart is where the design earns its keep, and where the sampling
bites. `aiVisibilityTrend.ts` is the best-reasoned file in the feature. It
splits time into the last 7, 28 or 90 days and the same span before it, and it
does not simply pool the answers. It builds cells, one per prompt, engine,
market and own-brand identity, computes each cell's rate, and averages cells.
Only cells with answers in both periods enter the comparison.

This is a paired design, and the right one. Adding ten new prompts this
week cannot move the comparison, because the new cells have no previous period.
Editing a prompt's wording starts a new prompt, so it does not either. Renaming
your brand starts new cells. Prompts that contain your brand name or domain are
excluded, so "Is Bitwarden safe to use?" does not prop up the rate. Failed and
empty collections never count as absent: `addAnswer` skips them, and the
comment above it says so. And if fewer than 90% of planned answers arrived in
either period (`MIN_COVERAGE = 0.9`), the rates still show but the change is
hidden.

What the cells cannot fix is that each one, per check, holds a single yes or
no. If the true chance that ChatGPT names you for a prompt is 50%, one check
tells you heads or tails. Pooling many prompts and many checks is the only way
to shrink that, and the arithmetic is unforgiving: the standard error of a
rate over $n$ independent yes/no answers is $\sqrt{p(1-p)/n}$, the difference
of two periods has $\sqrt{2}$ times that, and the smallest change a 95% test
separates from noise is about $1.96\sqrt{2}\,\sqrt{p(1-p)/n}$.

<NoiseBudget />

Take the launch post's own example, 20 prompts on ChatGPT and Gemini. Checked
weekly with the 7-day window, each period holds one check of 40 answers. At a
50% mention rate the per-period error is 7.9 points and the smallest
trustworthy week-over-week change is about 21.9 points. Daily checks put 280
answers in each 7-day period, and the band narrows to about 8.3 points; that
costs about \$2.40 a month on the hosted plan. Weekly checks read over the
28-day window get you 160 answers a period for \$0.32 a month and a band of
about 11 points, at the price of seeing a change a month late.

Two assumptions in that arithmetic pull opposite ways, and the widget's
caption says which. Treating every prompt as having the same rate is the worst
case for a given average; real prompts that are nearly always yes or always no
contribute less noise. Treating runs as independent is the best case. If
DataForSEO's scraper or the engine returns much the same answer on consecutive
days, the effective sample is smaller than the count. I have no way to measure
that correlation from outside, so I read the band as an order of magnitude.

The Explorer has a related trap. Rerunning the same prompt with the same
models within seven days is free, which the blog presents as a perk. In
`promptExplorer.ts` that is `PROMPT_RESPONSE_TTL_SECONDS = 7 * 24 * 60 * 60`:
the second run is the first run's answer from the cache. It is the right
billing choice, and it means the Explorer cannot show you variance at all.
If you want to see how much an answer moves, rephrase the prompt by one word.

Nothing here is a defect in OpenSEO. One answer per prompt and engine is what
\$0.002 buys, and the tool tells you so before every check. The trouble is the
dashboard: a column of percentages with no interval invites you to read a
5-point move as news. I would like to see the interval drawn on the trend
chart; the code already has every number needed to draw it.

## What \$0.002 per answer is made of

The blog's price is "\$0.002 per answer". The constant behind it is
`AI_RECORD_COST_USD = 0.0012` in `src/shared/ai-visibility.ts`, which matches
DataForSEO's published standard-queue price of \$0.0012 per LLM Scraper result.
For AI Overviews the comment splits it into a results page at \$0.0006 plus the
async overview load at \$0.0006, and DataForSEO refunds the second half when
Google shows no overview.

On the hosted service, `src/shared/billing.ts` applies `SEO_DATA_COST_MARKUP =
1.28`, converts to credits at 1,000 per dollar, and rounds up. One answer is
0.0012 × 1.28 = \$0.001536, which ceils to 2 credits, or \$0.002. This per-answer figure is the estimate, and the number the blog quotes. The settlement is different.
`settleUsageCredits` charges each provider call separately on DataForSEO's
reported cost, and one `task_post` carries a whole batch of up to 100 answers.
Twenty ChatGPT answers settle at 31 credits, not 40. At batch sizes like that
the hosted price lands at about \$0.0015 per answer, which is the README's
"28% extra for every request" and not the 67% that the per-answer estimate
implies. So the estimate is honest in the direction that matters: it
overstates, and the credit check before a run uses it, so a run never starts on
credits that cannot cover it.

The blog's examples check out against those rules. Twenty prompts on two
engines is 40 answers, \$0.08 per check estimated; weekly is about \$0.32 a
month and daily about \$2.40, both at the planning figures of 4 and 30 checks a
month in `scheduledChecksPerMonth`. Self-hosted, the same 40 answers cost
\$0.048 paid straight to DataForSEO. Prompt Research is the expensive call: a
`llm_mentions` request at \$0.1 plus \$0.001 per row, up to 100 rows, plus a
50-deep live results page at \$0.008. At the full 100 rows that is \$0.208
raw and about \$0.27 with the markup, close to the docs' "about \$0.25 per
keyword". It is cached for a day per organisation.

I did not check the "\$100+/month" for incumbent AI-visibility products. It is
plausible for the enterprise tools, but I have no primary source in front of
me, so I leave it as the launch's claim.

## Running it yourself

The licence is MIT, copyright Ben Senescu, with no extra terms. It is cleaner than most "open-source alternative" launches I have read; the
[treg](/articles/treg) teardown, for contrast, found an Apache 2.0 licence
with a clause forbidding hosting.

There are two self-hosting paths. Docker runs the Cloudflare Workers runtime
(`workerd`) inside a Node 22 image, keeps state in a volume, and runs with
`AUTH_MODE=local_noauth`, which means no login at all; the docs say to put it
behind your own authenticated proxy. A built-in scheduler looks for due checks
every five minutes, so scheduled tracking works with no host cron. The
recommended path is Cloudflare itself, which the README says works on the free
plan. Either way you need a DataForSEO API key, which is a paid account, so
`runsOn` for this page is honestly "an API", not "your laptop". Setup research
and generated prompt suggestions also want an `OPENROUTER_API_KEY`; the
default model is `openai/gpt-5.6-luna`, overridable with `OPENROUTER_MODEL`.

Two defaults are worth knowing before you run it. Telemetry is on: heartbeats
with aggregate counts tied to a random install ID, every five minutes for the
first two hours and then at most daily, documented as never including URLs,
keywords or prompts. `OPENSEO_TELEMETRY_DISABLED=1` turns it off. And tracking
cannot collect ChatGPT answers for one country in the shared market list,
location 2275 (Palestine), per `AI_UNSUPPORTED_LOCATIONS`.

The MCP server deserves a sentence. An agent can research, explore, configure
and read results, but a paid check needs a `maxCostUsd` the user approved from
a quote, and `aiVisibilityRuns.ts` refuses the run if the price has since
gone up. Letting an agent spend money is the risky part of agent tooling, and
this is a clean way to bound it. The [jev-linkmap](/articles/jev-linkmap)
piece is a reminder of how quickly an agent's tuning bill outgrows the
headline price when nothing bounds it.

## What this can and cannot see about a site like this one

Back to why I looked. Everything this site does for discoverability is on the
input side. `citationsFromBody()` in `lib/jsonld.tsx` puts up to 20 arXiv,
DOI, GitHub and Hugging Face links into each article's `citation` field.
`/llms.txt` is an index for agents, the `.md` twins are clean text, and
`robots.txt` declares `ai-train=yes, search=yes, ai-input=yes` for every
crawler. None of that can be observed from the outside except through what
the engines eventually say.

OpenSEO measures exactly that output, for three surfaces, for the questions
you choose. For a site like this, the owned-citation rate is the only number
that matters: nobody asks ChatGPT about "Satyajit Ghana" as a brand, but
someone might ask how SGLang's radix cache works, and the question is whether
the answer links here. The Citations tab would show which pages get cited
instead, which is a concrete thing to learn.

What it cannot do is connect an input to an output. It never sees whether
`llms.txt` was fetched, whether a page was in a training set, or which
retrieval index a scraper's session hit. It does not track Claude or
Perplexity at all, and it reads answers from whatever session state the
vendor's scraper has, not a logged-in user with memory and history. The
questions it suggests are People-also-ask questions, so they describe Google
searchers rather than chat users. And the noise band above is the ceiling on
attribution: if I ship a JSON-LD change and the owned-citation rate moves by 6
points on a weekly check, the tool has told me nothing. With daily checks over
a 28-day window, a move of that size starts to mean something, and even then
it says that the answers changed, not why.

Still, it is more than I have now. The thing I would actually do is pick a
dozen questions this site's articles answer better than most, track them on ChatGPT and AI Overviews daily (the two engines that, in the launch screenshot, cite the most different sources), and read the Citations tab monthly for which pages win
instead. The rate on the chart I would treat as a smoke alarm, not a
thermometer.

## How I checked

I shallow-cloned `every-app/open-seo` at `deb4491` (2026-10-06) and read the
AI-visibility path end to end: the workflow, the DataForSEO clients and price
table, the parser, the matcher, the trend and results services, the billing
helpers and the relevant UI components. The mention widget is a line-for-line
port of `markdownToText`, `aiMentionSpans` and `ownedAiDomain`, minus the NFD
offset mapping, and I ran the ported functions on its four presets in Node to
confirm the outcomes it shows. I did not run OpenSEO or call DataForSEO. The
Wilson intervals and the noise arithmetic are mine, computed from the counts
in the launch screenshots and from the formula stated above. DataForSEO's
\$0.0012 standard-queue price is from its LLM Scraper pricing page. The launch
posts and their thread were read through the fxtwitter mirror; the figures are
the launch card from X and the five screenshots committed with OpenSEO's own
launch post under `web/public/blog/ai-visibility-in-openseo/`.
