We measured 2,592 AI answers: which engines cite anyone at all
Perplexity has never once answered without citing a source. ChatGPT declines to cite anyone about one time in twenty. Both numbers come from our own runs, and here is the method.
8 min read · Published 2026-08-28
Almost everything written about AI visibility is downstream of somebody else's study. That is not a criticism, it is a description of a young field: the good posts cite Ahrefs or a Princeton preprint, attribute the figures properly, and add reasoning on top. We have written some of them.
What follows is different in one respect. Every number below was measured by us, on our own runs, and can be recomputed from the queries at the bottom. We publish it because the most useful thing we found is not something you can reason your way to, and because a company that sells measurement should be willing to publish a number it does not control.
What was measured
Between 13 July and 25 August 2026 we ran 186 distinct buyer-shaped prompts against four AI engines on behalf of 19 sites, and stored every answer verbatim along with the sources each engine returned. That produced 2,592 answers that came back successfully, 191 where the engine was unavailable, and 6 outright errors.
The unavailable runs matter and they are excluded from everything below. An engine that did not answer is not a measurement of absence, and counting it as a zero is the single most common way a visibility number gets quietly inflated in the wrong direction.
The finding: engines differ enormously in whether they cite at all
For each successful answer we asked one narrow question: did the engine cite any source at all, anyone, not necessarily the brand we were measuring. The spread was far wider than we expected.
| Engine | Answers measured | Cited nobody | Rate | Days sampled |
|---|---|---|---|---|
| Perplexity | 751 | 0 | 0.0% | 13 |
| Google AI Mode | 601 | 9 | 1.5% | 8 |
| Claude | 570 | 10 | 1.8% | 11 |
| ChatGPT | 670 | 122 | 18.2% | 10 |
Perplexity has never once, across 751 answers on 13 separate days, produced an answer without citing something. That is a structural fact about the product rather than a quirk of our sample: Perplexity's whole interface is built around numbered sources, and it appears to have no mode in which it answers bare.
ChatGPT is the opposite end and needs an honest caveat. The raw 18.2% is misleading, because 94 of those 122 uncited answers landed on a single day, 26 July, when 94 of 98 successful runs came back with no citations at all on the same model and the same model version as every other day. That is not a market event. Something in the retrieval path was not working, on their side or ours, and we cannot tell which.
Excluding that day, ChatGPT answered without citing anyone in 28 of 572 answers, or 4.9%. So the honest statement is: about one ChatGPT answer in twenty cites nobody, and one day in our window is unusable and is reported as unusable rather than averaged in.
Why we now exclude that day automatically
Finding it by hand was luck. The fix was not, and it is worth describing because it is the part most easily skipped.
A day where an engine cites nobody far more often than it normally does tells you nothing about whether a particular brand lost citations, so it must not enter a baseline. Our breakout detector compares a prompt against its own median citation rate over prior runs; if instrument zeros accumulate in that median, the median sags, and once they are half the history the median hits zero. At that point the detector reports that a prompt has never cited you before, about a prompt that cited you in every real run. It is not an inflated number, it is a false sentence.
So a run is now marked suspect and dropped from both the totals and the baselines when at least half of that engine's answers cited nobody, and that share is at least three times the same engine's own median across at least three other runs, and the run's 95% lower bound clears both. Across 19 sites that rule fires on exactly one incident, the 26 July ChatGPT runs, and on nothing else.
We do not claim to know whether it was an extraction fault on our side or an engine that stopped citing that day. We cannot tell those apart, so we exclude the run and say why rather than pick the flattering explanation.
The part that changes what you do: a zero has two meanings
Every tool in this category, ours included until this month, reports one number: the share of answers that cite you. That number collapses two situations that call for opposite work.
An answer can cite nobody, which means there is no incumbent and the job is to become the source the engine reaches for. Or it can cite other people and not you, which means there is an incumbent, by name, and the job is to displace a specific page. Told only that you appear in 12% of answers, you cannot tell which of those you are facing.
Splitting them on one site in our corpus, across 517 measured answers: 31 cited the site, 400 cited somebody else, and 86 cited nobody. Per engine, displacement ran between 74% and 86%. Almost every buyer question that site cares about already has an incumbent, which is a completely different brief from the one the headline percentage implies.
What this sample cannot tell you
Nineteen sites is not a market. They are the properties of one operator, weighted toward content publishing, and they are not a random sample of anything. A different vertical would very likely produce different displacement rates, and we would not be surprised to be wrong about the magnitude.
What the sample can support is narrower and, we think, more useful: the per-engine citation behaviour is a fact about the engines rather than about our sites, it is measured over hundreds of answers per engine across many days, and the direction of the differences is far too large to be sampling noise. Perplexity citing on 751 of 751 answers is not a close call.
Every share we publish inside the product carries a 95% Wilson interval for exactly this reason, and a change has to clear a significance test before we call it a change. The figures above are counts over a stated window, not projections, and we have not dressed them up as more than that.
How to check us
The queries are below. If you run an AI-visibility tool and your own corpus disagrees with ours, we would rather know: the useful version of this field is one where these numbers get argued about with data, not one where every vendor publishes a proprietary score and nobody can check anything.
You can also run the same instrument against your own domain for free at seoaio.ai/audit. It shows the whole result without an account, including the parts where it has nothing to report.
How this was measured
- Window
- 13 July 2026 to 25 August 2026, read on 28 August 2026.
- Sample
- 2,592 successful answers, 186 distinct prompts, 19 sites, 4 engines. 191 unavailable runs and 6 errors excluded. 32,326 stored citations, of which 834 pointed at the measured brand.
- Method
- One row per prompt run in prompt_runs, one row per returned source in citations, joined on prompt_run_id. An answer counts as citing nobody when its run has zero citation rows and its status is ok. Per-engine rates are counts over that engine's successful runs, not modelled.
- Limits
- One operator's portfolio, not a random sample; content-publishing heavy. The 26 July ChatGPT runs are excluded as suspect under the published rule and the raw figure is shown alongside so the exclusion is visible rather than silent.
Sources
- Aggarwal et al., GEO: Generative Engine Optimization (arXiv 2311.09735) The Princeton study our content rules cite for citation and statistics tactics.
- Wilson score interval The interval every share in our product carries, and why small samples read as less certain.
- RFC 9309, Robots Exclusion Protocol The standard behind the crawler-access half of AI visibility.
- The llms.txt convention The proposal our discovery checks lint against.
Related reading
See your own AI visibility measured, not guessed
One free crawl, a real score, and which named AI engines can currently recommend your site. No account needed.
