Methodology
What we measure, how we measure it, and — just as importantly — what this method cannot tell you.
How a check runs
- We generate seven questions a real customer might ask about a business like yours — two naming your business directly, the rest describing what you do and where.
- Each question is asked once on 5 assistants: ChatGPT, Claude, Perplexity, Gemini and Grok. All of them run with live web search enabled.
- Every answer is read by a classifier that decides whether your business was recommended, mentioned ambiguously, or not mentioned at all — and records which competitors appeared.
- Separately, we fetch your homepage and check the structural signals assistants rely on: schema.org markup,
llms.txt, your page title and meta description.
That produces 5 × 7 = 35 independent observations, each of which you can click to see the evidence behind the call.
Why we don't report a percentage
AI assistants do not return the same answer twice. Ask the same question an hour apart and the wording, the sources, and sometimes the recommendations change. That is inherent to how these models work, not a flaw in the tool.
So a single check is a sample, not a rate. Reporting “cited in 3% of queries” would imply a stable population figure we have not measured. We report what we actually observed: cited in N of 35 checks.
What this can establish is the extreme case. Being cited once or twice across 35 independent observations, spanning five different models and seven different phrasings, is strong evidence of genuine invisibility. You would need implausible luck to see that from a business assistants regularly recommend.
Why re-runs use the same questions
The first check on a domain establishes its question set, and every later check reuses it. If the questions changed each time, you would be comparing different measurements and calling the difference progress.
Even so, expect some movement between runs. A difference of one or two cells is normal variation. Treat a sustained shift across several checks as signal; treat a single cell as noise.
What this does not tell you
- How often real buyers see you. We measure whether assistants cite you for representative questions, not your share of actual conversations, which no one outside the AI companies can observe.
- Why a given answer came out that way.The models don't explain their reasoning, and we don't pretend to reconstruct it.
- What will happen if you change something. We can identify missing structural signals and content gaps. We cannot promise a specific outcome from fixing them.
- Anything about paid placement. This measures organic citation only.
When a platform fails
If an assistant errors or times out, that cell is marked as an error rather than counted as “not cited”. Treating a failed request as an absence would understate your visibility. Error cells are excluded from the citation count, and we monitor for platform outages so a broken column doesn't quietly distort reports.
Questions about the method
If something here doesn't hold up, we want to know — write to hello@citationchecks.com. A measurement product that won't explain its measurements isn't worth much.