Security controls affect automated access
Bot management, rate limiting and firewall rules can affect automated clients. That may be intended for scrapers and unintended for AI retrieval agents.
Your robots.txt is a declaration. Your edge decides what actually gets through. On a Cloudflare-fronted site those two can disagree without anyone noticing, in either direction.
Cloudflare sits in front of your origin and makes its own decisions about automated traffic. Those decisions are legitimate and often exactly what a security team intended. They are also independent of robots.txt, which is a published request rather than an enforcement mechanism.
So a site can invite an agent in robots.txt and still refuse it at the edge, or restrict it in robots.txt while the edge lets it through. Neither state is visible from the file alone, and neither is visible from a dashboard that only reads content.
This audit measures both sides and keeps them apart. That separation is the product: a declared policy is a fact about your file, an observed response is a fact about one request we made, and merging them produces a claim neither supports.
Bot management, rate limiting and firewall rules can affect automated clients. That may be intended for scrapers and unintended for AI retrieval agents.
An interstitial served with a 200 status is not the page. A client that cannot solve it receives markup that contains none of your content.
Apex to www, HTTP to HTTPS, or a country redirect. Each hop can change what an identified automated client receives.
A 403 or 429 to an automated request is a different fact from a 200 with a challenge page, and the remediation differs.
A restriction can live at the origin and be visible only as a response that passed through Cloudflare unchanged.
Rules added for one incident that were never reviewed again, still shaping automated access months later.
What can be established from outside, stated as observation rather than inference.
Your robots.txt graded across the frozen panel, with the exact line behind every verdict.
What our identified crawler received: status code, headers we retain, redirect chain and whether the body was the page or an interstitial.
The full chain as an automated client experiences it, including the URL finally reached.
Where the declared policy and the observed response disagree, stated as a difference rather than a diagnosis.
Findings that can only be confirmed inside your own Cloudflare account, described precisely enough for your team to check them.
What is present in the server-rendered HTML that reaches a non-rendering fetcher.
A clear statement of where declared policy and observed access diverge, with the evidence for each.
Exact robots.txt lines and the edge settings to review, written so your team can apply or assess them directly.
The specific things to confirm inside your own Cloudflare account, because some facts are only visible from there.
Ordered by likely commercial effect, with primary content restrictions ahead of secondary paths.
The Full AI Access Audit includes a free re-check after your fixes.
We do not publish evasion techniques or detection internals. The report tells you what to change on your own infrastructure.
Delivered remotely; no access to your Cloudflare account is required for the audit itself.
Access is necessary for retrieval and citation but does not guarantee either. We report what Digital Dominator’s identified audit crawler observed; that is evidence about our crawler and is never presented as proof of another provider’s crawler behaviour, and we do not request from a vendor network or impersonate a vendor crawler. Training crawlers, search and discovery crawlers, user-triggered fetchers and data-use controls are not interchangeable and are graded separately. llms.txt is optional and emerging, not a ranking requirement. One point deserves emphasis on Cloudflare-fronted sites: we do not state that a particular product feature caused a restriction unless that is what was observed. An edge may also answer two requesters differently, so one crawler’s result cannot establish another crawler’s treatment.
For the difference between a declared block and a real one, read how to check whether your site is blocking AI, and why the headline blocking statistic misleads. Pricing and the full method are on the central AI Access Audit page.