llms.txt: Does It Matter?
By Seth Black · Published
Every few weeks someone asks me whether they need an llms.txt file. Usually a tool flagged it as “missing,” or a consultant pitched it as the thing that gets you cited by ChatGPT. The honest answer is short: add one if you want, it takes an hour, and don’t expect it to change your rankings or your AI citations.
The longer answer is more interesting, because the evidence got a lot better this year.
What llms.txt actually is
llms.txt is a proposal from Jeremy Howard, first published in September 2024. The idea is simple. Websites are built for browsers, full of navigation, scripts, and layout. A language model working inside a limited context window would rather have a short, plain Markdown file that says what the site is and where the useful pages are.
It’s a proposal, not a standard. No standards body adopted it, and no crawler is required to read it. It sits at /llms.txt, next to robots.txt, but it does a completely different job. robots.txt tells bots what they’re allowed to fetch. llms.txt is a curated reading list.
The spec was revised in August 2026. Version 2 is based on two years of watching how people actually used it, and the changes page is worth a read.
The format
The spec is strict about order, but there isn’t much to it:
- An H1 with the name of the site or project. This is the only required part.
- A blockquote with a short summary.
- Optional paragraphs or lists with more detail. No headings here.
- Optional H2 sections, each containing a Markdown list of links. Each item is
[name](url), optionally followed by a colon and a note.
There’s a convention for an H2 called Optional that holds secondary links an agent can skip when it’s short on context. In v1 that section had mechanical meaning for tooling. In v2 it’s just a convention.
Version 2 also clarified a few things that matter if you run a large site:
- A file can live at a subpath like
/docs/llms.txt, and it covers the URLs under that path. When more than one file applies, agents are supposed to use the most specific one. - Pages can offer Markdown versions, either by appending
.mdto the URL or by replacing the extension. - Discovery uses standard link relations.
rel="alternate" type="text/markdown"points to a page’s Markdown version, andrel="describedby"points to the llms.txt that covers it. Either can go in an HTML<link>tag or an HTTPLinkheader.
What about llms-full.txt?
You’ll see llms-full.txt everywhere, and plenty of guides describe it as part of the spec. It isn’t. The llmstxt.org proposal never defines it.
llms-full.txt is a tooling convention: your whole documentation set concatenated into one big Markdown file. Mintlify says it built the format in collaboration with Anthropic, and docs platforms popularized it from there. The original proposal had its own expanded-context files generated by FastHTML tooling, and v2 dropped that tooling from the proposal entirely.
None of that makes llms-full.txt useless. A single file is convenient when you want to paste an entire docs site into a model, or when a coding agent wants everything in one request. Just don’t describe it as “the standard,” because it isn’t one.
Who actually reads it
This is the part people want to skip, and it’s the only part that matters.
Google Search: no. Google’s guide to optimizing for generative AI features says you don’t need to create “machine readable files, AI text files, markup, or Markdown” to appear in Search, including its AI features, and that creating them “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.” That’s about as clear as Google gets. John Mueller compared llms.txt to the keywords meta tag back in 2025, pointing out that none of the AI services had said they use it and server logs showed they didn’t even check for it.
The AI search crawlers: no public commitment. OpenAI, Anthropic, and Perplexity document their crawlers in detail. I cover them in the AI crawler robots.txt guide. None of those docs say their bots read llms.txt or treat it as a signal. OpenAI and Anthropic do publish llms.txt files for their own developer docs, and llmstxt.org points that out. Publishing one for your own docs doesn’t mean your crawler reads everyone else’s.
The log data: mostly nobody. Ahrefs published the best dataset I’ve seen, an analysis of 137,210 domains using server logs and live traffic from May 2026. Of the roughly 38,000 domains with a valid llms.txt, 97% got zero requests for that file all month. AI bots made zero requests for llms.txt files that didn’t exist, which means nothing is probing for them. Among the files that did get traffic, SEO audit tools were the biggest single category. GPTBot was the most common AI crawler fetching them, and Claude Code, a coding agent, showed up ahead of every AI search and assistant bot. Ahrefs notes its customer base skews technical, so treat the adoption numbers as an upper bound.
Chrome’s Lighthouse: sort of. Lighthouse added an agentic browsing category that looks for an llms.txt file. If you don’t have one, the audit is marked not applicable, because the Lighthouse docs say providing the file is optional “at the moment.” If you have one and it errors, it gets flagged. That’s a hygiene check, not a ranking factor.
Put that together and a pattern shows up. The things reading llms.txt are coding agents, developer tools, and people pointing a model at a docs site on purpose. Search crawlers building citation indexes are mostly not in that group.
Cost versus benefit
The cost is low. For most sites it’s one hand-written file. Some platforms, docs tools in particular, generate it automatically.
The benefit depends on who your visitors are:
- Developer docs, APIs, SDKs, open source projects. Real benefit. Programmers use coding agents that read documentation, and a clean llms.txt plus Markdown versions of your pages makes your docs easier to use inside those tools.
- SaaS and product sites. Small benefit. Someone asking an agent to compare tools might have it read your file. It’s cheap insurance and a decent forcing function for writing a clear one-paragraph description of what you sell.
- Local businesses, publishers, e-commerce catalogs. Close to zero today. Your effort is better spent on the pages themselves.
The real risk is opportunity cost. It’s easy to burn a week debating llms.txt while your product pages render blank without JavaScript. Most AI crawlers can’t see JavaScript-rendered content at all, and that problem costs you far more than a missing Markdown file.
One more thing. llms.txt is a statement of what you say your site is about, which is exactly why Mueller compared it to meta keywords. Anything it links to needs to match what’s on the actual pages. Stuffing it with claims your content doesn’t back up won’t fool anything that reads the pages too.
A minimal example
Here’s the shape of a sensible file for a small software product:
# Example Crawler
> Example Crawler is a desktop technical SEO crawler for Windows, macOS,
> and Linux. It audits sites for broken links, metadata, structured data,
> and JavaScript rendering problems.
Pricing is a one-time license. The documentation below covers
installation, the command-line interface, and report formats.
## Docs
- [Installation](https://www.example.com/docs/install.md): System requirements and setup
- [CLI reference](https://www.example.com/docs/cli.md): Every command and flag
- [Report formats](https://www.example.com/docs/reports.md): JSON, CSV, and HTML output
## Product
- [Pricing](https://www.example.com/pricing): License options
- [Free trial](https://www.example.com/free-trial): How the trial works
## Optional
- [Blog](https://www.example.com/blog): Articles on technical SEO
For a live example, our own file is at blackseoanalyzer.com/llms.txt, with a companion llms-full.txt.
Serve it as plain text with UTF-8, and add the discovery links from v2 if you’re publishing Markdown versions of your pages. In nginx:
location = /llms.txt {
default_type text/plain;
charset utf-8;
add_header Cache-Control "public, max-age=3600";
}
And in your page templates:
<link rel="alternate" type="text/markdown" href="/docs/install.md">
<link rel="describedby" href="/llms.txt">
How to keep it from going stale
A hand-written llms.txt rots the moment you rename a URL. Links that 404 in a file meant to help machines find your content defeat the entire purpose.
Generate it from the same source as your sitemap if you can. If you write it by hand, check its links whenever you check the rest of the site. A few lines of shell is enough:
curl -s https://www.example.com/llms.txt \
| grep -oE '\(https?://[^)]+\)' | tr -d '()' \
| while read -r url; do
printf '%s %s\n' "$(curl -s -o /dev/null -w '%{http_code}' "$url")" "$url"
done
Anything that isn’t a 200 needs fixing or removing.
My verdict
Add it. It’s cheap, it’s harmless, and for documentation-heavy sites it’s useful to the agents that already read it.
Don’t expect rankings, don’t expect AI citations, and don’t pay anyone who promises either one. The pages still do the work. Crawlable HTML, accurate structured data, fast responses, and content that answers real questions. That’s the same list it’s always been, which is why AIO and GEO keep turning out to be SEO.
If the spec picks up real support from the AI search crawlers, your file will already be sitting there. If it doesn’t, you lost an hour.
Before you polish your llms.txt, make sure the pages it points to are healthy. Try Black SEO Analyzer free for 14 days: every feature, no page limit, no credit card. Or see one-time pricing with no subscription.
-Sethers