What we measure, surface by surface
One page per thing we actually measure. On the engine side that is the AI answer engines we sample for citations. On the crawler side it is the AI crawler families we observe from your robots.txt and server logs. Each page states what we measure on that surface and, with equal weight, what we do not. There is no page here for a surface we do not measure.
The method behind every number lives on how we measure, and the line between the two lives on measured vs asserted.
Engines we sample for citations
What we measure about how Perplexity answers cite your site, how we sample it, and the limits of that number.
Starter plan and up
What we measure about whether Google's AI Overview cites your site, how the SERP is sampled, and why referral traffic is not part of the number.
Pro plan and up
What we measure about whether ChatGPT cites your site in its answers, how each sample is a real query, and the limits of that number.
Pro plan and up
Crawler families we observe
The OpenAI bots that can reach your pages, what each one is for, and how we observe their access and activity from your own robots.txt and server logs.
The Claude bots that can reach your pages, what each one does, and how we observe their access and activity. Claude is not one of our measured engines, and this page says so.
The Perplexity bots that can reach your pages, and how we observe their access and activity. Whether Perplexity then cites you is measured separately.
A robots.txt control, not a separate crawler, that governs whether Google may use content it already fetched to train and ground its generative products.
A robots.txt control on top of Applebot that governs whether Apple may use already-crawled content to train Apple Intelligence.
CCBot crawls the open web into a public dataset that many model builders train on, upstream of any single assistant.
