llms.txt is a Markdown file at the root of a website that gives AI assistants a map of it: the site's name, a short summary, then lists of links to the pages worth reading. Jeremy Howard proposed it in September 2024. Google has said Google Search ignores it, while other assistants do fetch it. We generate one for every site we run, as a cheap bet, and we don't promise anyone results from it.
What the format asks for
The proposal at llmstxt.org is short. The file is Markdown, and only one part is required, an H1 with the name of the project or site. After that, in order, come a blockquote with a short summary, any number of paragraphs or lists with more detail (no headings), and any number of H2 sections that each hold a list of links.
Each link line is [name](url), optionally followed by a colon and a note. An H2 called "Optional" is a convention for links an assistant can skip when it needs less context.
The proposal also recommends a clean Markdown version of each page at the same URL with .md added, so /pricing has /pricing.md. Pages announce that copy with a rel="alternate" link of type text/markdown, either as an HTML <link> element or as an HTTP Link header.
Who actually reads it
Be clear about this before spending a day on it. On 15 May 2026 Google published a guide to optimizing for its generative AI features, and it was blunt: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." It added that such files "will neither harm nor help" a site in Google Search.
Read the scope of that sentence. It's about Google Search. Assistants outside Google Search do fetch llms.txt, and an engine that reads a Next.js page raw gets half a screen of framework payload before it reaches the prose. A Markdown copy is simply easier for it to read.
So we treat llms.txt as a bet. It costs build time and nothing else, Google has said it carries no penalty, and nothing else on our sites depends on it. The AI Overview citations our own site picked up in September 2026 followed ordinary ranking work, which is what Google's guide says to expect.
What ours looks like
This is the top of the llms.txt that fabric.pro's build writes, trimmed:
# Fabric
> Fabric is an open-source AI agent platform that works from your
> company's own documents, code, tickets and conversations, and cites
> its source for every answer. ...
## Who builds Fabric
The company behind Fabric is a software company ...
## Product
- [Fabric: Open-Source AI Agent Platform for Every Team](https://fabric.pro/): the product page, with what it does, who it is for ...
- [Docs](https://app.fabric.pro/docs): setup, integrations, agents, the MCP gateway and self-hosting.
- [Source](https://github.com/Fabric-Pro): the open-source release ...
The summary in the blockquote is the same sentence the site uses in its meta description and on the home page. The detail block says who builds the product, with facts a person could check. Then come the sections, one link per page, each with a note that says what's on it.
Five things that make one useful
One description, everywhere. On 10 August 2026 we gave our consultancy's site one canonical description and made every other surface quote it, including llms.txt, the meta description, LinkedIn and directory listings. An assistant that finds three wordings of what a company does has three candidates to choose from.
Twins that are announced. A Markdown copy nobody can find doesn't help. On our consultancy's site the step that adds the Link headers ran before the hub pages' Markdown copies were written, so 21 hubs had a copy and no header pointing at it. We found that on 20 September 2026, and the hubs were the pages least able to afford it.
A pointer from robots.txt. Our llms.txt files had been live for months with nothing linking to them. robots.txt is the first file a crawler reads, and it has no registered field for llms.txt, so on 21 September we added a comment pointing at it. Parsers that follow RFC 9309 ignore comments, so the line costs nothing.
llms-full.txt beside it. That's the map followed by every page's Markdown copy, for an assistant that wants the whole site in one request. It isn't in the llmstxt.org proposal, and we generate it anyway because it falls out of the same build.
Generated by the build. A hand-written llms.txt goes stale the week someone adds a page. Ours is written from the built HTML after every export, and fabric.pro's build refuses to finish if a page's Markdown copy comes out too short or loses its questions section.
If you're choosing an llms.txt generator
About 480 people a month in the US search for one. Whatever you use, check that it reads your sitemap or your build output, takes each page's title and the paragraph that answers its question, groups pages under headings a person would recognise, and leaves out anything that redirects or carries noindex. If it also writes the Markdown copies and their headers, you've got the whole surface in one step.
Our generator, the Markdown copies and the check that keeps them in step with the pages are part of Heddle now, alongside the rest of the rules we use to make a site readable without a browser.
Questions
Does Google use llms.txt?
Google Search doesn't. Its May 2026 guide says Search ignores machine-readable AI text files and Markdown copies, and that they neither help nor harm a site there. Other assistants can and do fetch the file.
Where does llms.txt go?
At the root of the domain, as /llms.txt, served as plain text. Point to it with a comment in robots.txt so a crawler that reads robots.txt first will see it.
What's the difference between llms.txt and llms-full.txt?
llms.txt is a map, a summary followed by links to each page. llms-full.txt is that map followed by the full Markdown text of every page, so an assistant can read the whole site in one request. Only llms.txt is in the original proposal.
Should llms.txt replace sitemap.xml?
No. A sitemap tells search engines which URLs exist and when they changed. llms.txt tells an assistant what the site is and which pages are worth reading. A site needs both.