GEOGet cited by models

llms.txt and AI crawler policy

An AI crawler policy decides which language models may read your content and on what terms. It combines robots.txt directives for each AI user agent, an llms.txt file that points models at your best material, and a commercial position on training versus retrieval that most sites have never actually taken.

At a glance
Engagement
Five working days
Prerequisite
A decision on AI crawler access
First output
Published policy and robots rules
Common recommendation
robots.txt matters. llms.txt mostly does not

What llms.txt is, and what it is not

llms.txt is a markdown file at your domain root that points language models at your most useful content in a clean, readable form. It is a proposed convention, not a standard, and no engine is obliged to honour it. It is guidance, not access control.

The confusion worth clearing up first is that llms.txt and robots.txt do different jobs. robots.txt controls whether a crawler may fetch your pages at all — it is enforcement, honoured by every major crawler. llms.txt is a curation hint: here is the good stuff, in markdown, without the navigation and cookie banners. Treating one as a substitute for the other is the most common mistake we see.

Adoption is genuinely uneven, and anyone telling you otherwise is overselling. Some assistants and developer tools read llms.txt; the largest engines mostly do not, or do so inconsistently. That does not make it worthless — the cost of publishing one is an afternoon — but it does mean the file is not where the commercial decision lives.

The commercial decision lives in robots.txt, and it is genuinely two decisions that most sites have collapsed into one. Do you want your content used for training future models? And do you want it retrieved live so you can be cited today? Different user agents govern each, and they can be answered differently.

That distinction matters because the reflexive 2023 response — block everything with an AI in its name — answered the training question and silently answered the retrieval one too. Plenty of sites are invisible in ChatGPT and Perplexity today because of a decision made about training data, by someone who never intended to opt out of being recommended.

What you receive

You receive an audit of every AI user agent currently reaching or blocked from your site, a recommended policy separating training from retrieval, a published llms.txt curating your best content, corrected robots directives, and verification that the intended crawlers can actually fetch you.

Crawler access audit

Which AI user agents are reaching your content today, which are blocked, and which are blocked by a rule nobody remembers writing. Server logs where available, live fetch tests otherwise.

Training vs retrieval position

The two decisions separated and put to you plainly, with the commercial trade-offs of each. We recommend, you decide — this is a licensing question as much as a visibility one.

Corrected robots.txt

Per-agent directives that implement the position you chose, rather than a blanket rule that answers a question you were not asked.

Published llms.txt

A curated markdown index pointing models at your highest-value content, with the clean extracts that make it useful rather than a link dump.

Verification

Live fetch tests per agent after deployment, confirming the intended crawlers can reach the intended content and the excluded ones cannot.

Review schedule

New AI user agents appear regularly. You get a documented list and a quarterly review prompt, so the policy does not silently rot the way the last one did.

How the engagement runs

The engagement runs over five working days: audit which agents reach you today, put the training and retrieval decisions to you separately, implement the resulting directives, publish a curated llms.txt, then verify per agent that the policy does what you intended.

  1. 01

    Audit current access

    Every known AI user agent tested against your live site, cross-referenced with server logs where they exist. This regularly surfaces blocks nobody in the business knew were there.

    A per-agent access matrix

  2. 02

    Separate the two decisions

    Training use and live retrieval put to you as distinct choices, with the commercial consequences of each. Most clients end up allowing retrieval and restricting training, but it is your call to make.

    A written, signed-off policy position

  3. 03

    Implement and curate

    Robots directives written per agent, and an llms.txt built from your genuinely useful content rather than your sitemap. The curation is the part that takes judgement.

    Deployed robots.txt and llms.txt

  4. 04

    Verify and schedule review

    Fetch tests per agent confirming the policy behaves as intended, plus a documented quarterly review so new crawlers get a decision rather than a default.

    Verification report and review schedule

llms.txt compared with robots.txt

The two files are frequently conflated and do entirely different jobs. robots.txt is enforcement — it decides whether a crawler may fetch you at all, and every major crawler honours it. llms.txt is curation — it suggests what a model should read, and honouring it is optional.

Publishing llms.txt while robots.txt blocks AI agents achieves nothing. The robots file decides.
robots.txtllms.txt
PurposeControls whether a crawler may fetchSuggests which content is worth reading
StatusLong-established, universally honouredProposed convention, adoption uneven
EnforcementRespected by all major crawlersOptional — no engine is obliged
FormatDirectives per user agentMarkdown index with curated extracts
Gets you blockedYes — this is the file that excludes youNo — it cannot grant or deny access
Where the decision livesHereNot here

Signals you need this now

You need this when your robots file still carries a blanket AI block from 2023, when nobody can say whether GPTBot or PerplexityBot can reach you, or when a GEO baseline showed you absent from engines that would otherwise have every reason to cite you.

  • Your robots.txt blocks AI crawlers by a rule nobody remembers adding
  • A GEO baseline showed you absent from engines you should appear in
  • Nobody can say which AI agents can currently fetch your content
  • You have never separated the training question from the retrieval one
  • Legal or leadership want a defensible written position on AI use
  • You publish documentation or research worth curating for models
  • New AI crawlers appear and nobody decides anything about them

What changes afterwards

Access changes take effect within days, because crawlers re-read robots directives frequently. If a block was the constraint, citations can begin appearing almost immediately. If access was already fine, the honest outcome is a documented position and no visibility change at all.

Questions about llms.txt

Each answer is written to stand alone in 40 to 60 words — the shape an AI Overview or Perplexity citation lifts. Ships with FAQPage schema.

What is llms.txt?

A markdown file at your domain root that points language models at your most useful content in clean, readable form. It is a proposed convention rather than a standard, so engines may ignore it. It suggests what to read; it cannot control access.

Is llms.txt the same as robots.txt?

No. robots.txt decides whether a crawler may fetch your pages at all and is honoured universally. llms.txt suggests which content is worth reading and is honoured inconsistently. Publishing llms.txt while robots.txt blocks AI agents accomplishes nothing at all.

Should we block AI crawlers?

That is two questions. Blocking training use protects your content from future models. Blocking retrieval makes you invisible in today's answers. Most clients allow retrieval and restrict training, but it is a licensing decision we set out rather than make for you.

Does publishing llms.txt improve our visibility?

Modestly and inconsistently, because adoption is uneven. It costs an afternoon and helps with some assistants and developer tools. Anyone presenting it as the key to AI visibility is overselling — the robots directives matter far more.

How do we know our policy is actually working?

Live fetch tests per user agent after deployment, plus server log analysis where logs exist. We verify that intended crawlers reach the intended content and excluded ones do not, rather than assuming the file was written correctly.

How often does this need revisiting?

Quarterly is sensible, because new AI user agents appear regularly and an undecided crawler falls to whatever your default rule says. You get a documented agent list and a review prompt so the policy does not silently go stale.

Find out whether this is your constraint.

Thirty minutes with a senior strategist. We pull your live visibility while we talk and tell you plainly whether a llms.txt is what you need — or whether your problem sits somewhere else.

Book a discovery call →