Policy and infrastructure disagree
robots.txt declares one thing; the CDN, WAF or reverse proxy does another. Both are working as configured, and the combination was never reviewed as one system.
A fixed-price diagnostic for Heads of Digital, CMOs, SEO leads and platform owners who need to know what AI crawlers actually meet at the edge — and hand engineering a change list the same week.
On a large site, nobody owns crawler access end to end. Marketing owns the content, an agency may own robots.txt, platform engineering owns the CDN, and security owns the WAF. Each team makes reasonable decisions inside its own boundary, and the combined behaviour is something no one specified and no one is watching.
That is why declared policy and observed behaviour drift apart on enterprise estates more than anywhere else. A robots.txt can invite an agent that a bot-management rule quietly challenges. A redirect chain added for a locale launch can change what a non-rendering fetcher receives. None of it appears in a dashboard, because it is not a content problem.
This audit is scoped to answer that question and stop. It is a fixed-price diagnostic with a defined deliverable, not another platform to buy, integrate and renew, and not a project to reinvent internally.
robots.txt declares one thing; the CDN, WAF or reverse proxy does another. Both are working as configured, and the combination was never reviewed as one system.
The person who can change robots.txt is not the person who can change an edge rule. Findings without a named layer and an exact change stall between teams.
Apex to www, HTTP to HTTPS, country or language redirects. Each hop is a chance for a non-rendering fetcher to receive something other than the page.
Template robots files and platform defaults carried through migrations, still making decisions nobody made deliberately.
Access restrictions intended for a preview environment that were promoted along with everything else.
Training crawlers, search and discovery crawlers, user-triggered fetchers and data-use controls restricted together, when the commercial consequences differ sharply.
Two separate evidence classes, never merged, so a reader can always tell what was declared from what was observed.
Your robots.txt graded across a frozen panel of crawler, fetcher and control identities, with the exact source line behind every verdict.
What Digital Dominator’s identified audit crawler received: status, redirects followed, response characteristics and timing.
Where a difference appears between declared policy and observed response — CDN, WAF, reverse proxy or origin — reported as what was observed at each hop.
Whether a restriction touches primary public content or only secondary and operational paths. The commercial difference between the two is large.
The chain as an identified automated client experiences it, including where a hop changes the response.
Structured data, canonical hygiene and what appears in server-rendered HTML, because that is what a non-rendering fetcher receives.
One page a CMO or Head of Digital can act on, stating what is restricted, where, and what it plausibly affects.
Exact robots.txt lines, header and status-code changes, and edge settings, each tied to the evidence it rests on and the layer that owns it.
Findings that belong to the WAF or bot-management layer, described so a security reviewer can assess them without a marketing translation.
Ordered by likely commercial effect, not by count of findings, so the first three items are the ones worth doing.
The Full AI Access Audit includes a free re-check after your fixes, so the change is verified rather than assumed.
Portfolios, multi-brand estates and agency engagements are quoted per domain with consolidated reporting. Enquire with the domain list.
Enterprise, agency and multi-domain engagements are quoted on the domain list.
Access is necessary for retrieval and citation but does not guarantee either. We report what Digital Dominator’s identified audit crawler observed; that is evidence about our crawler and is never presented as proof of another provider’s crawler behaviour, and we do not request from a vendor network or impersonate a vendor crawler. Training crawlers, search and discovery crawlers, user-triggered fetchers and data-use controls are not interchangeable and are graded separately. llms.txt is optional and emerging, not a ranking requirement. On an enterprise estate one further limit matters: a single observation is a point in time from one vantage. An edge can answer two requesters differently, so we record what was requested, what came back and when, rather than generalising from it.
For the wider programme this sits inside, read the enterprise AI visibility audit guide. For the underlying method, read how an AI access audit actually works. For the handoff itself, writing a developer brief that gets acted on covers how these change lists are structured. The central AI Access Audit page has pricing and the full method.