A widely repeated claim says that around 40% of websites are “blocking AI”. The figure is not fabricated, but it combines very different types of restriction into one headline number. The independently published PTODA C01 analysis shows why that headline tells the wrong story.
C01 measures declared AI crawler policy and the scope of restrictions visible in that policy. It does not establish how a particular CDN, WAF or origin responds to live crawler requests, and it does not measure AI rankings, recommendations, citations, visibility outcomes or commercial performance.
What PTODA measured
The independently published PTODA C01 AI Crawler Access Study analysed 2,699 sampled rows across five national cohorts—Australia, the United States, Great Britain, Singapore and India—representing 2,652 distinct domains on one frozen evidence base. The research is published through the independent PTODA research programme. Digital Dominator applies related crawler-access analysis commercially but does not own the research findings.
The number is technically true and practically misleading
Here is the trick of it. If you count a website as "blocking AI" the moment its robots.txt disallows any AI crawler from anything, you get the famous number: in the PTODA cohorts, the result varies by market. The statistic is not fabricated.
But look at what is being blocked and the picture inverts. Most of that blocking never touches primary public content generally relevant to AI retrieval systems. It targets carts, admin paths, search results pages, operational corners. When PTODA classified restriction by what the policy actually covers:
Restriction reaching primary public content remained a minority position across the five cohorts, ranging from 6.7 to 11.7 percent of policy-observed domains.
Whole-site exclusion of AI retrieval crawlers—the clearest “we block retrieval” policy position—ranged from 3.5 to 5.0 percent of policy-observed domains.
The gap between an “any restriction” headline and 3.5–5.0% whole-site exclusion of AI retrieval crawlers is not a rounding error. The two measures answer different questions.
Different restriction scopes hidden by one number
The second finding matters even more for individual businesses. Some restrictive policies carry signs of inheritance rather than a clearly documented decision: template robots files listing crawlers by the dozen, legacy rules and configurations created for a different problem. C01 records the published policy; it does not infer the organisation’s intent.
There is also a separate layer robots.txt cannot establish: enforcement. Firewalls and CDN bot protection act independently of the declared policy. Separate Digital Dominator enforcement testing records what its identified audit crawler receives; a 403 response, managed challenge, redirect or timeout is reported as a separate observation and is never attributed to C01.
Why the correction matters
If you rely on the 40 percent headline, every restriction appears equivalent. The corrected picture is more useful: whole-site exclusion of AI retrieval crawlers is uncommon, while narrower restrictions vary in scope and potential relevance. C01 does not infer intent. For an individual business, the practical question is therefore: what does our policy restrict, and does it match our intended access position?
That question has an answer, per site, with evidence. The ten-minute manual version is in our guide to checking your own robots.txt; the crawler-by-crawler mechanics are in the field guide to all 21 AI crawlers; and the full evidence-grade diagnosis - policy, enforcement and readability - is the AI Access Audit.
The research itself, including the frozen methodology and per-market findings, is published through the Periodic Table of Digital Authority research programme - the framework work we covered in this article.