GUIDE · AI SEARCH · OCTOBER 2026

What is llms.txt? A full example, how to create one, and who actually reads it.

llms.txt is a small Markdown file at the root of a site that hands AI tools a curated map of your most useful pages. In one 300,000-domain study, about one site in ten had one. The part most guides skip: Google says Search ignores it, and the crawlers behind AI answers almost never fetch it. This guide covers what the file is, who does read it, a complete example, and how to write yours in under an hour.

BY THE SEO AGENT TEAMUPDATED 2026-10-0214 MIN READ
Editorial cover image for a guide to llms.txt: what it is, an example file and who reads it
THE SHORT ANSWER

What is llms.txt, and does any AI engine read it?

llms.txt is a Markdown file served at yoursite.com/llms.txt that gives AI tools a short, curated list of your most useful pages, each with a one-line note. Jeremy Howard proposed it in September 2024 and revised it in August 2026. It controls nothing: robots.txt still decides which crawlers get in, and sitemap.xml still lists every URL for search engines. As of October 2, 2026, no operator of ChatGPT, Claude, Gemini or Perplexity has said its crawlers read the file, Google says Search ignores it, and in Ahrefs' May 2026 logs across 137,000 domains, 97% of valid llms.txt files got no requests at all. The readers it does have are coding agents such as Claude Code, which look for documentation. Create one if you publish docs or an API. For a content site it is an optional one-hour task with no measured effect on AI citations.

1. What is llms.txt?

llms.txt is a text file, written in Markdown, that sits at the root of a website (yoursite.com/llms.txt) and tells a large language model where the useful material on the site lives. It is not a list of every URL. It is a short, hand-picked index: the name of the site, a one-paragraph summary, and groups of links, each with a note on what the page contains. The idea is that an AI tool answering a question about your product can read one small file and go straight to the right pages, instead of wading through navigation menus, cookie banners and JavaScript.

Jeremy Howard, co-founder of the research lab Answer.AI, proposed the format in September 2024 and published it on its own site, llmstxt.org. It is a proposal, not a standard from a standards body, and no AI company is bound by it. Version 2, published in August 2026 after what the proposal calls two years of adoption, kept the file format the same and added ways for tools to find clean Markdown copies of your pages.

Two details from the proposal explain most of the confusion around it. First, the file is meant for inference, the moment a model is answering someone, not for training. Second, it is designed to pair with Markdown copies of your pages, so a model can read your content without the HTML around it. Neither detail makes any crawler fetch the file. Whether anything does is a separate question, and it decides whether the file is worth your time. It is also a much narrower job than LLM SEO, which is about getting your pages cited inside the answer itself.

2. Does any AI engine actually read llms.txt?

As of October 2, 2026, the honest answer is: the answer engines do not say they do, and the server logs say they mostly do not. Here is the evidence, newest first where it matters.

  • Google says no, in writing. Gary Illyes said at Search Central Live in July 2025 that Google does not support llms.txt and has no plans to. Google's guide to generative AI features in Search, published May 15, 2026, says you do not need “new machine readable files, AI text files, markup, or Markdown” to appear in AI search. A June 2026 addition says keeping llms.txt files “won't harm (nor help)” your visibility, because Google Search ignores them.
  • Almost nobody fetches it. Ahrefs analyzed May 2026 server logs across 137,000 domains. About 38,000 had a valid llms.txt, and only about 1,100 of those received a single request, so 97% were never fetched in the month. AI retrieval bots, the kind that fetch pages to answer a question, made about 1% of the requests. SEO audit tools made 21%.
  • One site, watched for 90 days. OtterlyAI published an llms.txt on a test site and logged its AI bot traffic: 84 of more than 62,100 AI bot requests went to the file, about 0.1%. The average content page on the same site got about 265.
  • No link to citations. SE Ranking checked about 300,000 domains in November 2025. 10.13% had an llms.txt, and having one showed no relationship with how often a domain was cited by AI models. Their prediction model got more accurate when the llms.txt variable was removed.
  • No commitment from the AI companies. We found no public statement from the companies behind ChatGPT, Claude, Gemini or Perplexity saying their crawlers (GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot) look for llms.txt on the open web. Their own developer docs do publish one: Claude's developer docs serve an llms.txt that links a Markdown copy of every docs page, and the developer docs for ChatGPT's API serve one that points to a separate llms.txt per documentation set.
Editorial illustration: a street mailbox with one card left sticking out of its slot, standing for an llms.txt file that is rarely collected

So who does read it? Coding agents. In the Ahrefs data, coding agents made about 10% of llms.txt requests, more than any other AI category, and Claude Code and GPTBot were the two most active individual bots. When a developer asks an agent to integrate an API, the agent goes looking for documentation, and a clean index of Markdown pages is exactly what it wants. That is the reader John Mueller of Google had in mind in May 2026 when he said “for non-developer sites, I don't think this makes much sense, even with more agentic traffic.”

One more reader arrived in May 2026. Chrome's Lighthouse 13.3 added an Agentic Browsing category that checks for llms.txt. A missing file is marked not applicable, not failed. A server error fails, and so does a file whose links are not written as Markdown links. The audit measures whether the file parses, not whether anyone uses it, and it sits outside the SEO category. If your real goal is to show up when someone asks ChatGPT, Perplexity or Google's AI Overviews a question, the levers are the ones in our generative engine optimization guide: pages crawlers can reach, a direct answer near the top, and facts worth quoting.

3. llms.txt vs robots.txt vs sitemap.xml.

Three files sit at the root of a site, and they do three different jobs. Mixing them up is how people end up believing llms.txt can block AI training. It cannot. Only one of the three has any force.

FileJobWho reads itEnforces anything?
robots.txtSays which crawlers may fetch which pathsGooglebot, GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBotYes, for crawlers that honor it
sitemap.xmlLists every URL you want indexed, with last-modified datesSearch engine crawlersNo, it is a discovery hint
llms.txtPoints AI tools to a short list of your best pages, in MarkdownCoding agents, some AI tools, audit tools, LighthouseNo

robots.txt is the one with teeth, and the one that decides your AI visibility. The big AI companies each run separate bots for separate jobs. GPTBot and ClaudeBot collect pages that may be used to train ChatGPT and Claude models. OAI-SearchBot and Claude-SearchBot index pages for ChatGPT search and Claude search. ChatGPT-User, Claude-User and Perplexity-User fetch a page when a person asks about it, and Perplexity-User generally ignores robots.txt for that reason. Google-Extended is not a crawler at all: it is a robots.txt token that controls whether Google may use your pages for its Gemini models, and it does not affect Google Search.

That split lets you opt out of training and stay in AI search. This robots.txt does exactly that:

# Out of AI training, still visible in AI search
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Notice what is not in it: llms.txt. The two files never reference each other. A crawler that honors robots.txt checks it before it fetches any URL, llms.txt included, so if you block a bot there, listing pages in llms.txt does not let it back in. sitemap.xml stays what it always was, the complete list of URLs for search engines, and it should still exist. If you are auditing all three, start with the crawl checks in our technical SEO audit checklist, because a robots.txt mistake costs you far more than a missing llms.txt ever will.

4. The llms.txt format, line by line.

The format is plain Markdown with a fixed order. Only the first line is required. Everything else is there to make the file useful to a model with a limited context window.

  1. An H1 with the site or project name. The only required section.
  2. A blockquote summary. One line starting with >, holding the few sentences a model needs to understand the rest of the file.
  3. Optional detail, with no headings. Paragraphs or bullet lists with the facts a model should not get wrong: prices, limits, what you do not do.
  4. H2 sections of links. Each is a Markdown list in which every item is a link, optionally followed by a colon and a note: - [Title](URL): note.
  5. An Optional section. By convention, links an agent can skip when it needs a shorter context.
# Site or project name

> One or two sentences: what this is, who it is for.

Facts a model should not get wrong (plain paragraphs or a list, no headings).

## Section name

- [Page title](https://example.com/page.md): what is on the page

## Optional

- [Less important page](https://example.com/other.md): skip when short on context

Version 2 of the proposal, published in August 2026, added three things. Markdown copies of pages can now live at the page URL with .md appended (page.html.md) or with the extension replaced (page.md). Each HTML page can point to its Markdown copy with a link tag using rel=“alternate” and type=“text/markdown”, and to the llms.txt that covers it with rel=“describedby”, either in the page head or as an HTTP Link header. And files can live in subfolders: /docs/llms.txt covers the pages under /docs, and the most specific file wins. Version 2 also dropped the companion tool that expanded the file into one big context, so the Optional section is now only a naming convention.

Editorial illustration: a card-file index box with four divider tabs rising above rows of cards, standing for the grouped link sections of an llms.txt file

You will also see llms-full.txt. It is not in the proposal. It is a convention many documentation platforms adopted: one file that holds the full text of every page the llms.txt links to, so a tool can load a whole documentation set in one request. Claude's developer docs link one at the bottom of their llms.txt. It makes sense for docs that fit in a model's context window, and very little for a blog.

One practical rule from Lighthouse: write every link as a real Markdown link. A file of lines like “Pricing: /pricing” fails the Agentic Browsing audit with “File does not appear to contain any links”, even when every URL works. The same discipline helps any parser that does read the file, which is the reason for the strict format in the first place.

5. How to create llms.txt in six steps.

A first version takes under an hour for most sites. The time goes into choosing pages and writing the notes, not into the syntax.

  1. Pick 10 to 30 pages. Start from the questions an AI tool most often gets wrong about you: what you cost, what you do and do not do, how the product works, your policies. Add your documentation and your two or three strongest guides. Leave the full URL list to sitemap.xml; a reading list with 400 links is not a reading list.
  2. Write the H1 and the summary. The name, then a one or two sentence blockquote, then the hard facts as a short list. Prices, plan limits and the things you do not do belong here, because they are what a model is most likely to state wrongly.
  3. Group the links under H2 sections. Docs, Product, Guides and Optional cover most sites. Write each note as a plain description of what is on the page (“plans, limits and what counts as an invoice”), not marketing copy.
  4. Decide on Markdown copies. If you publish documentation, serve a clean .md copy of each page and link to those; most documentation platforms can generate them. For a marketing site, linking to your normal HTML pages is fine. Google's John Mueller has argued against building separate Markdown pages just for bots, and nothing in the log data suggests he is wrong for a non-documentation site.
  5. Publish it at the root. The file must load at yoursite.com/llms.txt with a 200 status and a plain-text content type. On a static or Next.js site, drop it in the public folder. On WordPress, Yoast SEO and Rank Math both added an llms.txt setting in 2025 that builds the file from your pages and skips anything marked noindex; check what it picked, because automatic selection favors recent posts over your key pages.
  6. Test it and keep it current. Open the URL in a browser, then run Lighthouse from Chrome DevTools and check the Agentic Browsing result. Regenerate the file on every deploy or whenever prices, docs or key pages change. A link to a page that now returns 404 is worse than no link.

If your customers connect AI tools to your product rather than just reading about it, llms.txt is the read-only first step. The heavier version gives an agent live tools instead of a reading list, and our SEO MCP server guide shows what that looks like in practice.

6. When llms.txt is worth an hour, and when it is not.

The file costs almost nothing to keep, and Google says it does no harm. The real cost is the belief that it does something it does not. Here is how to decide.

  • Worth it: docs, APIs and developer tools. Coding agents are the one group that measurably reads the file, and a developer wiring up your API is a customer in the middle of adopting you.
  • Worth it: when it is automatic. If your documentation platform or SEO plugin already generates it, turn it on, check the output, and move on.
  • Cheap insurance: SaaS with facts models get wrong. Prices, limits and “we do not do that” facts in the summary cost nothing and help any tool that does look.
  • Not worth more than an hour: content sites, local businesses, blogs. Expect no measurable effect on AI citations or Google rankings, which is what the 300,000-domain study found.

What does move AI citations is less exotic. Let the search crawlers in (OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot). Serve pages whose text is in the HTML, not assembled by JavaScript after load. Put the direct answer in the first two sentences under a heading that matches the question, which is the core of answer engine optimization. Publish facts with a date and a source, because a model that cites you is quoting a claim. Then measure whether it is working with one of the AI visibility tools that track citations, rather than guessing.

If you want the whole picture of how ChatGPT, Perplexity, Gemini and Google's AI features choose sources, our AI search optimization page lays out each engine, and the guide to getting cited in AI answers covers what models retrieve and why. None of it starts with a text file.

WORKED EXAMPLE

A worked example: a complete llms.txt file.

Take a fictional six-person company, Ledgerline, that sells an invoicing API and web app to agencies. Two problems pushed them to write an llms.txt. Developers kept asking coding agents to integrate the API and got answers built on an old version. And when prospects asked ChatGPT about Ledgerline, it sometimes said the product ran payroll, which it does not. The company, its URLs (on the reserved .example domain) and its logs are invented. The format and the pattern of who fetched the file match the studies above.

Editorial illustration: a closed hardback book with three bookmark ribbons hanging out, standing for a short reading list into a larger site

Here is the whole file they published at ledgerline.example/llms.txt, written in about 50 minutes:

# Ledgerline

> Ledgerline is an invoicing API and web app for agencies that bill by the hour. It turns tracked time into invoices, collects card and bank payments, and syncs paid invoices to accounting software.

Key facts:

- Plans: Starter is $29 a month for up to 50 invoices. Studio is $79 a month for unlimited invoices and 5 seats. Card fees are passed through at cost.
- The REST API authenticates with an API key sent as a Bearer token. Versions are set by date in the Ledgerline-Version header.
- Ledgerline does not run payroll, file taxes or issue cards.

## Docs

- [Quickstart](https://ledgerline.example/docs/quickstart.md): Create an API key and send a first invoice
- [Invoices API](https://ledgerline.example/docs/api/invoices.md): Create, update, send and void invoices
- [Payments API](https://ledgerline.example/docs/api/payments.md): Card and bank payments, refunds, payout timing
- [Webhooks](https://ledgerline.example/docs/webhooks.md): Event list, retry schedule, signature checks
- [Errors](https://ledgerline.example/docs/errors.md): Every error code with its fix

## Product

- [Pricing](https://ledgerline.example/pricing.md): Plans, limits and what counts as an invoice
- [Integrations](https://ledgerline.example/integrations.md): Accounting and time-tracking apps it syncs with
- [Security](https://ledgerline.example/security.md): Data location, encryption, compliance reports
- [Does Ledgerline do payroll?](https://ledgerline.example/help/payroll.md): No, and which payroll tools it exports to

## Guides

- [Billing retainers](https://ledgerline.example/guides/retainers.md): Recurring invoices with unused hours rolled over
- [Late payment reminders](https://ledgerline.example/guides/reminders.md): The default schedule and how to change it

## Optional

- [Changelog](https://ledgerline.example/changelog.md): API changes by date
- [About](https://ledgerline.example/about.md): Founders, location and contact

What is working in it. The summary says what the product is in one sentence a model can repeat. The key facts carry the exact things models got wrong: the prices, the API's authentication and versioning, and a flat “does not run payroll”. Every link points at a Markdown copy and carries a note that describes the page instead of selling it. The payroll question has its own page, linked by its question, so a tool scanning the file finds the answer by name. The changelog and about page sit under Optional because a tool can skip them. Thirteen links in total, not the 240 URLs in their sitemap.

DAYS 1-30 · WHO FETCHED /LLMS.TXT (INVENTED LOGS)

41 requests. 17 from SEO audit tools checking whether the file exists, 12 from coding agents (Claude Code most often) while developers integrated the API, 9 from scanners and unidentified bots, 3 from the team testing it. None from GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot.

WHAT FIXED THE DEVELOPER PROBLEM

The file did. Agents that read it went to the current Markdown docs, and the support tickets about deprecated endpoints dropped.

WHAT FIXED THE PAYROLL ANSWER

Not the file. The help page “Does Ledgerline do payroll?”, with the answer in its first sentence, did the work once OAI-SearchBot and Claude-SearchBot had crawled it, because answer engines quote pages, not index files.

That split is the whole lesson. llms.txt solved the problem whose reader exists, developers working through an agent, and did nothing for the problem whose reader does not. The wrong answer in AI search was fixed the way it always is: a crawlable page that answers the exact question, stated plainly, near the top. Ledgerline now has a backlog of thirty more of those questions, which is where the real hours go.

Common mistakes with llms.txt.

  1. Expecting it to get you cited. It is sold as an AI SEO shortcut because it is easy to ship and easy to audit. The citation studies found no effect, and Google ignores it. Ship it if it is cheap, then put the effort into pages, which is the real difference covered in how GEO differs from classic SEO.
  2. Using it to block AI training. The file has no directives and no crawler treats it as a rule. Opting out of training happens in robots.txt, bot by bot (GPTBot, ClaudeBot, Google-Extended).
  3. Pasting the whole sitemap into it. Generators make this tempting. A file with hundreds of links gives a model no idea which pages matter, which defeats the point of a curated index.
  4. Writing links as plain text. “Pricing: /pricing” reads fine to a person and fails Lighthouse's check. Use the Markdown link form for every entry.
  5. Listing pages bots cannot reach. Pages blocked in robots.txt, behind a login, or rendered only by JavaScript stay unreadable no matter what the file says. Check them the way an assistant would: ask Claude with web search or ChatGPT with search on to summarize a listed page and see what comes back.
  6. Writing it once and forgetting it. The summary holds your prices and limits. When those change and the file does not, any tool that reads it now repeats the old facts with your name on them.
BEFORE YOU GO

llms.txt is a one-hour file, not a strategy.

Write it if you publish docs, turn it on if your platform generates it, and keep it short and current. Then spend the time where AI engines actually look: crawlable pages, the right bots allowed in robots.txt, answers stated plainly at the top, and facts with a source. That is the ongoing work, and it is the part most teams cannot keep up week after week.

That daily loop is the job we built our AI SEO agent to do. The SEO Agent turns it into fixed stages:

  • Real keyword data. Live monthly search volume and difficulty behind every article it plans, checked against the pages you already have so it never writes a duplicate.
  • Answer-first, fact-checked drafts. Every claim gets a cited source, and unsupported sentences are rewritten or removed.
  • A quality gate. A draft that fails is refused, not published.
  • Native publishing, daily. Into WordPress, Webflow, Shopify, Wix, Ghost or Framer, or anywhere else through a webhook, with every stage laid out on the automated SEO content pipeline page.

It costs a flat $99 a month after a free trial, you cancel in one click from inside the app, and the story of why we built it is a short read.

QUESTIONS

Common questions about llms.txt.

Missing something? Ask us directly.

What is llms.txt?

llms.txt is a plain-text file written in Markdown and served at the root of a website (yoursite.com/llms.txt). It gives AI tools a short, curated index of the site: a title, a one-paragraph summary, and groups of links to the most useful pages, each with a short note. Jeremy Howard proposed it in September 2024 and published version 2 in August 2026. It is a proposal, not a standard, and it controls nothing: robots.txt still decides which crawlers may fetch your pages.

Do ChatGPT, Claude, Gemini or Perplexity read llms.txt?

Not in any way their operators have announced. As of October 2, 2026 we found no public statement from the companies behind ChatGPT, Claude, Gemini or Perplexity saying their crawlers look for llms.txt on the open web, and the log data agrees: in Ahrefs' May 2026 study of 137,000 domains, AI retrieval bots made about 1% of llms.txt requests. Coding agents are the exception. They made about 10% of requests, and Claude Code was one of the two most active bots.

Does Google use llms.txt?

No. Google's Gary Illyes said in July 2025 that Google does not support it and has no plans to. Google's guide to generative AI features in Search, published May 15, 2026, says you do not need new machine readable files, AI text files, markup or Markdown to appear in AI search, and a June 2026 addition says keeping llms.txt files will neither harm nor help your visibility in Google Search because Google Search ignores them.

Does llms.txt help SEO or AI visibility?

There is no evidence that it does. SE Ranking checked about 300,000 domains in November 2025 and found that having an llms.txt had no measurable effect on how often a domain was cited by AI models. Google Search ignores the file. The things that do move AI citations are crawlable pages, search crawlers allowed in robots.txt, direct answers near the top of a page, and facts with a source.

What is the difference between llms.txt and robots.txt?

robots.txt controls access: it tells crawlers such as Googlebot, GPTBot, ClaudeBot and PerplexityBot which paths they may fetch, and well-behaved crawlers obey it. llms.txt controls nothing. It is a reading list that suggests which pages an AI tool should look at. If robots.txt blocks a crawler, listing a page in llms.txt does not reopen it.

Can llms.txt stop AI companies from training on my content?

No. To keep pages out of AI training, disallow the training crawlers in robots.txt: GPTBot for ChatGPT models, ClaudeBot for Claude models, and the Google-Extended token for Gemini. Leave the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) allowed if you still want to appear in AI search answers. Some user-triggered fetchers, such as Perplexity-User, generally ignore robots.txt because a person asked for the page.

Where do I put llms.txt?

At the root of your domain, so it loads at https://yoursite.com/llms.txt, returned with a 200 status and a plain-text content type. Version 2 of the proposal also allows files in a subfolder, such as /docs/llms.txt, which cover the pages under that path; when two files apply, the more specific one wins.

What is llms-full.txt?

A single file that contains the full text of every page an llms.txt links to, so an AI tool can load a whole documentation set in one request. It is not part of the llms.txt proposal; it is a convention that documentation platforms adopted. Claude's developer docs, for example, link one from the bottom of their llms.txt. It makes sense for documentation that fits in a model's context window and very little sense for a blog.

How do I create an llms.txt file in WordPress?

Yoast SEO and Rank Math both added an llms.txt setting in 2025. Turn it on and the plugin builds the file from your pages and skips URLs marked noindex. Open the result at yoursite.com/llms.txt and edit the selection if your SEO plugin allows it: an automatic file tends to list recent posts rather than the 10 to 30 pages that answer questions about your business.

Should my site have an llms.txt?

If you publish developer documentation or an API, yes: coding agents are the one group that measurably reads the file, and many documentation platforms generate it for you. For a content site, local business or blog, it is an optional one-hour task with no measured payoff in AI citations or Google rankings. Do it if it is cheap, and spend the real effort on the pages it points to.

PAST THE TEXT FILE

One researched, fact-checked article a day, written to be the page AI engines quote.

The SEO Agent picks keywords from live search volume and difficulty, puts the answer first, cites a source for every claim, refuses drafts that fail its quality gate, and publishes natively to WordPress, Webflow, Shopify, Wix, Ghost or Framer. $99 a month after a free trial. Cancel in one click.

FREE TRIAL · CANCEL IN ONE CLICK