robots-observatory

Measuring whether AI crawlers actually respect robots.txt.

Certain paths on this host are disallowed for one specific AI crawler each. Requests that arrive are recorded. A hit counts as a violation only if the client address appears in the IP ranges that vendor publishes itself; a request carrying a crawler's User-Agent from anywhere else is counted separately as spoofing, never as a violation.

Results are aggregated once a day and committed to a public git repository, so every figure has a history and every change to it is a diff.

Scorecard

Observation window 2026-08-21 → 2026-08-29 (2 days with data, 7 recorded hits, 1 site(s)).

BotVendorAttributionHitsVerifiedViolationsSpoofedIn graceNot disallowedOther UAFirst seenLast seen
GPTBotopenaiexclusive10000012026-08-212026-08-21
OAI-SearchBotopenaiexclusive0000000——
ChatGPT-Useropenaiexclusive0000000——
ClaudeBotanthropicambiguous60n/a00062026-08-292026-08-29
Claude-Useranthropicambiguous00n/a0000——
PerplexityBotperplexityexclusive0000000——
Perplexity-Userperplexityexclusive0000000——
CCBotcommoncrawlexclusive0000000——
DuckAssistBotduckduckgoexclusive0000000——
MistralAI-Usermistralexclusive0000000——
Google-Extendedgoogleambiguous00n/a0000——
Applebot-Extendedappleambiguous00n/a0000——

Reading this table. Verified counts hits whose client address was found in the vendor's own published ranges. Violations counts verified hits on a path that robots.txt disallowed for that bot at the time of the hit, outside the 48h grace period that follows any robots.txt change. Spoofed counts requests that claimed the bot's User-Agent from an address the vendor does not publish; those are never counted as violations, and say nothing about the vendor. Attribution is n/a in the violations column for bots whose vendor publishes one range list covering several crawlers, because a hit cannot be attributed to one of them.

Excluded from violation counts for that reason: Applebot-Extended, Claude-User, ClaudeBot, Google-Extended.

Caveats