Crawler access audit
Which AI user agents are reaching your content today, which are blocked, and which are blocked by a rule nobody remembers writing. Server logs where available, live fetch tests otherwise.
An AI crawler policy decides which language models may read your content and on what terms. It combines robots.txt directives for each AI user agent, an llms.txt file that points models at your best material, and a commercial position on training versus retrieval that most sites have never actually taken.
Typically a fixed-scope one-off, delivered in five working days.
llms.txt is a markdown file at your domain root that points language models at your most useful content in a clean, readable form. It is a proposed convention, not a standard, and no engine is obliged to honour it. It is guidance, not access control.
The confusion worth clearing up first is that llms.txt and robots.txt do different jobs. robots.txt controls whether a crawler may fetch your pages at all — it is enforcement, honoured by every major crawler. llms.txt is a curation hint: here is the good stuff, in markdown, without the navigation and cookie banners. Treating one as a substitute for the other is the most common mistake we see.
Adoption is genuinely uneven, and anyone telling you otherwise is overselling. Some assistants and developer tools read llms.txt; the largest engines mostly do not, or do so inconsistently. That does not make it worthless — the cost of publishing one is an afternoon — but it does mean the file is not where the commercial decision lives.
The commercial decision lives in robots.txt, and it is genuinely two decisions that most sites have collapsed into one. Do you want your content used for training future models? And do you want it retrieved live so you can be cited today? Different user agents govern each, and they can be answered differently.
That distinction matters because the reflexive 2023 response — block everything with an AI in its name — answered the training question and silently answered the retrieval one too. Plenty of sites are invisible in ChatGPT and Perplexity today because of a decision made about training data, by someone who never intended to opt out of being recommended.
You receive an audit of every AI user agent currently reaching or blocked from your site, a recommended policy separating training from retrieval, a published llms.txt curating your best content, corrected robots directives, and verification that the intended crawlers can actually fetch you.
Which AI user agents are reaching your content today, which are blocked, and which are blocked by a rule nobody remembers writing. Server logs where available, live fetch tests otherwise.
The two decisions separated and put to you plainly, with the commercial trade-offs of each. We recommend, you decide — this is a licensing question as much as a visibility one.
Per-agent directives that implement the position you chose, rather than a blanket rule that answers a question you were not asked.
A curated markdown index pointing models at your highest-value content, with the clean extracts that make it useful rather than a link dump.
Live fetch tests per agent after deployment, confirming the intended crawlers can reach the intended content and the excluded ones cannot.
New AI user agents appear regularly. You get a documented list and a quarterly review prompt, so the policy does not silently rot the way the last one did.
The engagement runs over five working days: audit which agents reach you today, put the training and retrieval decisions to you separately, implement the resulting directives, publish a curated llms.txt, then verify per agent that the policy does what you intended.
Every known AI user agent tested against your live site, cross-referenced with server logs where they exist. This regularly surfaces blocks nobody in the business knew were there.
→ A per-agent access matrix
Training use and live retrieval put to you as distinct choices, with the commercial consequences of each. Most clients end up allowing retrieval and restricting training, but it is your call to make.
→ A written, signed-off policy position
Robots directives written per agent, and an llms.txt built from your genuinely useful content rather than your sitemap. The curation is the part that takes judgement.
→ Deployed robots.txt and llms.txt
Fetch tests per agent confirming the policy behaves as intended, plus a documented quarterly review so new crawlers get a decision rather than a default.
→ Verification report and review schedule
The two files are frequently conflated and do entirely different jobs. robots.txt is enforcement — it decides whether a crawler may fetch you at all, and every major crawler honours it. llms.txt is curation — it suggests what a model should read, and honouring it is optional.
| robots.txt | llms.txt | |
|---|---|---|
| Purpose | Controls whether a crawler may fetch | Suggests which content is worth reading |
| Status | Long-established, universally honoured | Proposed convention, adoption uneven |
| Enforcement | Respected by all major crawlers | Optional — no engine is obliged |
| Format | Directives per user agent | Markdown index with curated extracts |
| Gets you blocked | Yes — this is the file that excludes you | No — it cannot grant or deny access |
| Where the decision lives | Here | Not here |
You need this when your robots file still carries a blanket AI block from 2023, when nobody can say whether GPTBot or PerplexityBot can reach you, or when a GEO baseline showed you absent from engines that would otherwise have every reason to cite you.
Access changes take effect within days, because crawlers re-read robots directives frequently. If a block was the constraint, citations can begin appearing almost immediately. If access was already fine, the honest outcome is a documented position and no visibility change at all.
Each answer is written to stand alone in 40 to 60 words — the shape an AI Overview or Perplexity citation lifts. Ships with FAQPage schema.
A markdown file at your domain root that points language models at your most useful content in clean, readable form. It is a proposed convention rather than a standard, so engines may ignore it. It suggests what to read; it cannot control access.
No. robots.txt decides whether a crawler may fetch your pages at all and is honoured universally. llms.txt suggests which content is worth reading and is honoured inconsistently. Publishing llms.txt while robots.txt blocks AI agents accomplishes nothing at all.
That is two questions. Blocking training use protects your content from future models. Blocking retrieval makes you invisible in today's answers. Most clients allow retrieval and restrict training, but it is a licensing decision we set out rather than make for you.
Modestly and inconsistently, because adoption is uneven. It costs an afternoon and helps with some assistants and developer tools. Anyone presenting it as the key to AI visibility is overselling — the robots directives matter far more.
Live fetch tests per user agent after deployment, plus server log analysis where logs exist. We verify that intended crawlers reach the intended content and excluded ones do not, rather than assuming the file was written correctly.
Quarterly is sensible, because new AI user agents appear regularly and an undecided crawler falls to whatever your default rule says. You get a documented agent list and a review prompt so the policy does not silently go stale.
Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you plainly whether a llms.txt is what you need — or whether your problem sits somewhere else.