Before you pay anyone anything, you can answer a surprisingly large part of the question yourself, in about ten minutes, with nothing but a browser. This is the exact manual check we recommend to every business owner - what to look at, what the lines mean, and where the manual method genuinely runs out of road.

Step one: read your own robots.txt

Go to yoursite.com/robots.txt. Everyone has one of these, almost nobody has read theirs. It is a plain text file, and it is your published policy toward every automated visitor on the web.

You are looking for two kinds of lines. User-agent: lines name a crawler. Disallow: lines beneath them say what that crawler may not fetch. Two patterns matter:

Pattern one - the named block. Something like User-agent: GPTBot followed by Disallow: /. That is a whole-site block of a specific AI crawler. Check the name against the big five for AI visibility: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot. If any of those five sits above a Disallow: /, AI systems in that family cannot cite you. Names like GPTBot, ClaudeBot and CCBot are training crawlers - blocking those is a content-use choice, not a visibility problem. Our field guide to all 21 AI crawlers tells you which is which.

Pattern two - the wildcard. User-agent: * with Disallow: lines applies to everyone, AI included. A few disallowed paths like /cart or /admin is normal housekeeping. Disallow: / under the wildcard means your site tells every polite crawler on earth to leave.

Step two: ask the AI systems directly

Open ChatGPT, Perplexity, and Google (for AI Overviews) and ask each what your business does, or ask for recommendations in your category and location. You are checking two things: are you mentioned at all, and when a system cites sources, are you ever one of them? Do it in a private window so your own history does not flatter the result. If competitors are cited for questions you should own, you have a visibility gap - and step one may have just told you why.

Step three: check the years-old template problem

Look at the whole robots file again and ask one question: does this read like something anyone at your business deliberately decided? Long lists of crawler names, legacy SEO-tool blocks and contradictory rules can indicate an inherited template. The independently published PTODA C01 study provides the wider research context; Digital Dominator applies research-informed analysis to determine what the evidence means for your individual site.

Where the manual check runs out

Here is the honest limit. Everything above tests your policy - what your site says. It cannot test your enforcement - what your firewall and CDN actually do when an AI crawler connects. Bot protection acts before robots.txt is ever consulted, and the two disagree constantly: we routinely find sites whose policy says welcome while their firewall serves AI crawlers a challenge page or a 403. You cannot see that from a browser, because to your browser the site works perfectly.

The manual check also cannot tell you whether your pages are readable once fetched - whether your structured data, entity markup and content structure let a machine understand who you are well enough to cite you.

The evidence-grade version

That enforcement-and-readability layer is exactly what our AI Access Audit tests: all 21 crawler identities graded at the policy layer, live identified requests recording what your infrastructure actually does, a readability review, and a fix list with the exact lines to change. We wrote up how the audit works here. But start with the ten-minute version above - a decent share of businesses find something worth fixing before ever spending a dollar.