In practice
The most common and most costly misunderstanding is treating it as access control. It grants nothing and blocks nothing. A site that blocks AI crawlers in robots.txt and publishes an llms.txt has locked the door and posted directions to it — a combination we see more often than it should exist.
Access decisions live in robots.txt, which is an established standard that major crawlers including AI crawlers respect. Whether to allow those crawlers is a genuine strategic choice: blocking them removes you from the answers your buyers read, which for most businesses is a worse outcome than the training exposure it prevents.
As a tidy index of your best content, llms.txt costs little and does no harm. As a visibility strategy it is a bet on a convention that may not take hold, and there is no reliable public evidence that major models fetch it. We publish that assessment plainly because clients are being sold otherwise.
Not to be confused with
- robots.txt
- robots.txt is an established standard that controls crawler access and is honoured. llms.txt is a proposal that controls nothing.
Questions
Does llms.txt actually work?
There is no reliable public evidence that major models fetch it, and no engine has committed to honouring it. It is cheap and harmless to publish; expecting measurable visibility from it is not supported.
Can llms.txt stop models training on our content?
No. It is not an access-control mechanism and grants no permissions. Restricting crawler access is done in robots.txt, and even that governs crawling rather than every possible use.
Should we publish one anyway?
It is low cost and does no harm, so if maintaining it is genuinely easy, publish it. Just do not treat it as a substitute for the crawler policy decision in robots.txt.
