Tracked AI crawler family
Google-Extended
A robots.txt control, not a separate crawler, that governs whether Google may use content it already fetched to train and ground its generative products.
Feeds: Gemini and AI Overviews training.
The bots in this family
Google-Extended does not fetch anything on its own. It is a robots.txt token that tells Google whether content Googlebot already crawled may be used to train and ground Gemini and AI Overviews. Blocking it does not remove you from Google Search; it withholds you from that generative use.
| Bot | robots.txt token | Purpose |
|---|---|---|
| Google-Extended | google-extended | Training |
How we observe it
- Access. We fetch your robots.txt and evaluate it per named token using RFC 9309 matching. Each bot reads allowed, blocked, or unknown. If robots.txt could not be fetched, every row is unknown and nothing is guessed.
- Activity. This family is a robots.txt control, not a fetching agent, so it sends no distinct requests to observe. We report its access directive and stop there rather than imply a server-hit reconciliation that does not exist.
The same catalog powers the free crawler check tool and the Analyze crawler panels.
What we do not measure here
Stating the limit is the point. A number is only trustworthy next to the boundary of what it does not cover.
- Server-hit activity. Google-Extended is a policy token, not a fetching agent, so there are no distinct requests to reconcile. We observe the robots.txt directive, and nothing more is claimed.
- Whether Google honored the directive internally, which is not observable from outside Google.
- Citation in AI Overviews, which is measured on its own engine page rather than inferred from this control.
Access is a precondition for being cited, not the citation itself. Whether Gemini and AI Overviews training then cites you is measured separately on the Gemini and AI Overviews training engine page.
Every surface we measure
Engines we sample
Crawler families we observe
