llms.txt vs robots.txt — what each file actually controls

robots.txt is an established standard that major crawlers, including AI crawlers, respect for access control. llms.txt is a proposed convention for pointing models at your key content, with limited adoption and no guarantee of being read. Use robots.txt for policy; treat llms.txt as low-cost positioning.

The two options

robots.txt

The long-standing file controlling which crawlers may access which paths.

Choose this when
  • You need to allow or block specific AI crawlers
  • Sections of the site should not be crawled at all
  • You want a decision that crawlers actually honour

llms.txt & Crawler Policy

llms.txt

A proposed file pointing language models at a site's most useful content in a clean format.

Choose this when
  • Your documentation is large and worth signposting
  • The cost of maintaining it is genuinely low
  • You accept it may never be read

llms.txt & Crawler Policy

How they differ

 robots.txtllms.txt
StandingEstablished, widely honouredProposed, limited adoption
Controls accessYesNo
Honoured by major crawlersYesNot reliably
Risk of getting it wrongHigh — can deindex a siteLow
EffortLow, but needs careLow to moderate, ongoing

What actually decides it

The decision that actually matters is in robots.txt: whether to allow AI crawlers. Blocking them removes you from the answers your buyers read, and for most businesses that is a worse outcome than the training-data exposure it prevents. It is a strategic choice rather than a technical one, and it should be made deliberately rather than copied from a template.

llms.txt is cheap enough to be worth doing and oversold enough to be worth being sceptical about. There is no reliable evidence that major models fetch it, and no engine has committed to honouring it. As a tidy index of your best content it costs little; as a visibility strategy it is a bet on a convention that may not take hold.

The error worth avoiding is treating llms.txt as access control. It grants nothing and blocks nothing. A site that blocks AI crawlers in robots.txt and publishes an llms.txt has simply locked the door and posted directions to it.

Questions

Should we block AI crawlers?

For most businesses, no. Blocking them removes you from assistant answers your buyers are actively reading. Publishers whose product is the content itself have a genuine case; most other businesses do not.

Does llms.txt actually work?

There is no reliable public evidence that major models fetch it, and no engine has committed to honouring it. It is cheap and harmless to publish; expecting measurable visibility from it is not supported.

Can llms.txt stop models training on our content?

No. It is not an access-control mechanism and grants no permissions. Restricting crawler access is done in robots.txt, and even then it governs crawling rather than every possible use.

Related services

Still not sure?

Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you which of these two is actually binding for you — including when the answer is neither. Book a discovery call →