A growing share of your customers ask an AI assistant before they ever open a search results page. ChatGPT, Claude, Perplexity and Google's AI features answer those questions by reading and citing a small number of websites they can access and trust. Which raises a question most business owners have never been able to answer with evidence: can those systems actually see your site?

That is the question the AI Access Audit exists to answer. Not with a score, not with a hunch - with logged observations. This article explains what the commercial audit does, how research-informed evidence discipline shapes it, and why apparent “AI blocking” must be interpreted carefully.

Where the audit came from

The independently published PTODA C01 AI Crawler Access Study analysed 2,699 sampled rows across five national cohorts—Australia, the United States, Great Britain, Singapore and India—representing 2,652 distinct domains under one frozen evidence base. The research is published through the independent PTODA research programme. Digital Dominator applies related crawler-access analysis commercially but does not own the research findings.

The widely quoted statistic says something like "40% of websites block AI." Measured properly, that number falls apart. Roughly two in five sites do block at least one AI crawler somewhere in their robots.txt - but most of that blocking never touches the primary public content AI systems actually read. Genuine whole-site exclusion sat at around 4 to 5 percent. A large share of the rest was inherited: firewall defaults, copy-pasted robots templates, security rules installed years ago for a different problem. Blocking that nobody decided.

Which means the useful question for any individual business is not "do websites block AI?" It is: which case are you? Deliberate policy, inherited accident, or open but unreadable. The audit answers that for one site: yours.

What we actually check

The audit works in three layers, because access can fail at three different levels - and they fail independently.

Layer one: what your site says (policy)

Your robots.txt is your published policy toward machines. We capture it, hash it, and grade it across the configured commercial crawler panel - separating retrieval/search agents such as OAI-SearchBot, Claude-SearchBot and PerplexityBot from training controls such as GPTBot and ClaudeBot. These are genuinely different decisions. Blocking training is a legitimate business choice. Blocking retrieval can prevent direct retrieval and citation from your own site. Plenty of sites do the second while intending the first.

We also check for llms.txt and other machine guidance: whether you publish it, whether it is valid, and whether it agrees with your actual policy - because in our research, they often did not.

Layer two: what your site does (enforcement)

Here is the part almost nobody tests. Your robots.txt can say "welcome" while your firewall says "go away." CDN bot protection, security plugins, and WAF rules act on requests before your policy is ever consulted. In the Full Audit we make live, honestly identified requests and record exactly what comes back: a clean response, a challenge page, a 403, or a silent timeout. Policy and enforcement disagree far more often than you would expect - and when they do, the firewall wins and your policy is a fiction.

Layer three: what your site means (readability)

Access is necessary but not sufficient. Once an AI system can fetch your pages, it still has to understand them: who you are, what you do, whether you are the same entity it has seen elsewhere. So the Full Audit reviews your structured data, entity markup, canonical hygiene and how much of your content is actually present in the HTML a machine receives. A reachable site that reads as noise still does not get cited.

What you receive

Every audit produces a verdict board with crawler-policy findings backed by the actual robots.txt lines, plus a separate live-retrieval result for Digital Dominator's identified audit crawler backed by the observed status code, headers and challenge evidence. Then the part that actually matters: a prioritised fix list with the exact changes written out. The precise robots.txt lines. The specific header. The CDN setting. Ready to hand to whoever maintains your site, most fixes taking minutes.

The report also tells you what not to change - because over-correction is real, and a well-meaning developer "fixing" something that was already right is how sites end up worse off.

The Full Audit closes the loop with a free re-check: apply the fixes, tell us, and we verify each one worked with fresh observations.

What the audit is not

It is not a citation guarantee. Nobody can promise you AI citations, and anyone who does is selling something else. Access is the floor, not the ceiling: it determines whether you are eligible to be read and cited. What the audit gives you is certainty about that floor, with evidence - which, in our experience measuring thousands of sites, is exactly the thing most businesses have never once verified.

Where to start

If you just want to poke around, the free DD Diagnostic Suite covers the basics - robots.txt, schema, and a quick AI scan. When you want the full evidence-grade answer for your own site, the AI Access Audit starts at $495, takes five business days, and ends with a fix list rather than a lecture.