The two options
The long-standing file controlling which crawlers may access which paths.
- You need to allow or block specific AI crawlers
- Sections of the site should not be crawled at all
- You want a decision that crawlers actually honour
A proposed file pointing language models at a site's most useful content in a clean format.
- Your documentation is large and worth signposting
- The cost of maintaining it is genuinely low
- You accept it may never be read
How they differ
| robots.txt | llms.txt | |
|---|---|---|
| Standing | Established, widely honoured | Proposed, limited adoption |
| Controls access | Yes | No |
| Honoured by major crawlers | Yes | Not reliably |
| Risk of getting it wrong | High — can deindex a site | Low |
| Effort | Low, but needs care | Low to moderate, ongoing |
What actually decides it
The decision that actually matters is in robots.txt: whether to allow AI crawlers. Blocking them removes you from the answers your buyers read, and for most businesses that is a worse outcome than the training-data exposure it prevents. It is a strategic choice rather than a technical one, and it should be made deliberately rather than copied from a template.
llms.txt is cheap enough to be worth doing and oversold enough to be worth being sceptical about. There is no reliable evidence that major models fetch it, and no engine has committed to honouring it. As a tidy index of your best content it costs little; as a visibility strategy it is a bet on a convention that may not take hold.
The error worth avoiding is treating llms.txt as access control. It grants nothing and blocks nothing. A site that blocks AI crawlers in robots.txt and publishes an llms.txt has simply locked the door and posted directions to it.
Questions
Should we block AI crawlers?
For most businesses, no. Blocking them removes you from assistant answers your buyers are actively reading. Publishers whose product is the content itself have a genuine case; most other businesses do not.
Does llms.txt actually work?
There is no reliable public evidence that major models fetch it, and no engine has committed to honouring it. It is cheap and harmless to publish; expecting measurable visibility from it is not supported.
Can llms.txt stop models training on our content?
No. It is not an access-control mechanism and grants no permissions. Restricting crawler access is done in robots.txt, and even then it governs crawling rather than every possible use.
Related services
- llms.txt & Crawler Policy — GEO
- Technical SEO Audit — SEO
- Retrieval-Optimized Content — GEO
- GEO Baseline Audit — GEO
Still not sure?
Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you which of these two is actually binding for you — including when the answer is neither. Book a discovery call →
