Cloudflare-fronted · Edge behaviour · Remote delivery

Cloudflare AI Crawler Access Audit.

Your robots.txt is a declaration. Your edge decides what actually gets through. On a Cloudflare-fronted site those two can disagree without anyone noticing, in either direction.

A Cloudflare AI crawler access audit reports two separate things: what your robots.txt declares for each identity in a frozen 21-identity panel, and what Digital Dominator’s identified audit crawler actually received when it requested your site. Where the two differ, we record the observed response — status code, redirect chain or challenge — without asserting which specific product feature produced it unless that is what we observed.
Is your edge quietly answering AI crawlers differently from your robots.txt?
Fixed price. AUD $495 / USD $349 AI Access Check, or the AUD $1,195 / USD $799 Full AI Access Audit for deeper CDN, WAF, TLS and redirect analysis and a re-check after your fixes.
Book your AI Access Audit ›

01robots.txt and edge enforcement are separate systems

Cloudflare sits in front of your origin and makes its own decisions about automated traffic. Those decisions are legitimate and often exactly what a security team intended. They are also independent of robots.txt, which is a published request rather than an enforcement mechanism.

So a site can invite an agent in robots.txt and still refuse it at the edge, or restrict it in robots.txt while the edge lets it through. Neither state is visible from the file alone, and neither is visible from a dashboard that only reads content.

This audit measures both sides and keeps them apart. That separation is the product: a declared policy is a fact about your file, an observed response is a fact about one request we made, and merging them produces a claim neither supports.

02What can go wrong

Security controls affect automated access

Bot management, rate limiting and firewall rules can affect automated clients. That may be intended for scrapers and unintended for AI retrieval agents.

Challenge responses instead of pages

An interstitial served with a 200 status is not the page. A client that cannot solve it receives markup that contains none of your content.

Redirect chains at the edge

Apex to www, HTTP to HTTPS, or a country redirect. Each hop can change what an identified automated client receives.

Status codes that read as refusal

A 403 or 429 to an automated request is a different fact from a 200 with a challenge page, and the remediation differs.

Origin rules behind the edge

A restriction can live at the origin and be visible only as a response that passed through Cloudflare unchanged.

Cache and configuration drift

Rules added for one incident that were never reviewed again, still shaping automated access months later.

03What the audit examines

What can be established from outside, stated as observation rather than inference.

Declared policy, 21 identities

Your robots.txt graded across the frozen panel, with the exact line behind every verdict.

Observed edge response

What our identified crawler received: status code, headers we retain, redirect chain and whether the body was the page or an interstitial.

Redirect behaviour

The full chain as an automated client experiences it, including the URL finally reached.

Policy versus observation

Where the declared policy and the observed response disagree, stated as a difference rather than a diagnosis.

What needs your dashboard

Findings that can only be confirmed inside your own Cloudflare account, described precisely enough for your team to check them.

Machine readability

What is present in the server-rendered HTML that reaches a non-rendering fetcher.

04What you receive

The observed difference

A clear statement of where declared policy and observed access diverge, with the evidence for each.

A developer-ready change list

Exact robots.txt lines and the edge settings to review, written so your team can apply or assess them directly.

A dashboard checklist

The specific things to confirm inside your own Cloudflare account, because some facts are only visible from there.

Prioritised remediation

Ordered by likely commercial effect, with primary content restrictions ahead of secondary paths.

A re-check

The Full AI Access Audit includes a free re-check after your fixes.

No bypass advice

We do not publish evasion techniques or detection internals. The report tells you what to change on your own infrastructure.

05Fixed-price tiers

AI Access Check
AUD $495
includes Australian GST · USD $349
  • Executive access verdict
  • Frozen 21-identity declared-policy matrix
  • Exact robots.txt evidence per finding
  • Separate Digital Dominator crawler access observation
  • Prioritised developer-ready remediation
  • Human-reviewed before delivery
Book AI Access Check
Fix & Verify
Quoted
  • We implement the agreed changes
  • Bounded scope, quoted from your audit
  • Re-tested and verified after changes
  • Separate from any ongoing retainer
Enquire

Delivered remotely; no access to your Cloudflare account is required for the audit itself.

06What this evidence does not show

Access is necessary for retrieval and citation but does not guarantee either. We report what Digital Dominator’s identified audit crawler observed; that is evidence about our crawler and is never presented as proof of another provider’s crawler behaviour, and we do not request from a vendor network or impersonate a vendor crawler. Training crawlers, search and discovery crawlers, user-triggered fetchers and data-use controls are not interchangeable and are graded separately. llms.txt is optional and emerging, not a ranking requirement. One point deserves emphasis on Cloudflare-fronted sites: we do not state that a particular product feature caused a restriction unless that is what was observed. An edge may also answer two requesters differently, so one crawler’s result cannot establish another crawler’s treatment.

07Common questions

Can Cloudflare block AI crawlers even though robots.txt allows them?
Yes, and the reverse is also possible. robots.txt is a declaration and the edge is an enforcement layer; they are independent. That is why declared policy and observed access are reported as separate evidence classes and never inferred from one another.
Will you tell us which Cloudflare setting caused it?
Only if that is what we observed. From outside we can establish the response an identified crawler received; some causes can only be confirmed inside your own account, and those are listed as checks for your team rather than asserted as findings.
Do you need access to our Cloudflare dashboard?
No. The audit is performed externally with an identified crawler. Where a finding can only be confirmed from inside the account, the report says so and tells your team exactly what to look at.
Does a challenge page mean AI systems are blocked?
It means our identified crawler received an interstitial rather than the page on that request. It is evidence about our crawler at that moment, not proof of how any vendor crawler is treated, and we do not present it as such.

08Related reading

For the difference between a declared block and a real one, read how to check whether your site is blocking AI, and why the headline blocking statistic misleads. Pricing and the full method are on the central AI Access Audit page.

Is your edge quietly answering AI crawlers differently from your robots.txt?
Start with the central AI Access Audit, or book directly below.
Book your AI Access Audit ›
AI Access Audit — by city: Brisbane· Sydney· Melbourne· Perth· Canberra· Gold Coast | SEO Byron Bay