Measurements we ran ourselves and industry claims we traced to their sources: site migrations, redesigns, AI crawlers, and everything else that quietly changes the web. Written for SEO professionals who are tired of statistics nobody can footnote.
How we measure: the full method behind the AI-crawler dataset — sample, parser, denominators, limitations and a dated correction log — is at Methodology. It is open for comment, including the objections that make our numbers worse.
A robots.txt line declaring whether you allow AI training is about a year old. We found it on 1,104 of the web's most-visited domains, and found that 78.8% of them carry one byte-identical string a CDN wrote for them. Where a person wrote the line, 42.7% allow training.
We read robots.txt on the web's most-visited domains and found that sites block model-training crawlers several times more often than the AI-search crawlers that send readers back. Then we found that most of that gap was never a decision.
The most-quoted statistic in migration SEO, traced back through every citation we could find. What the evidence actually supports, and the honest numbers to use instead.