
What Is LLMs.txt? Exploring Its Function and How to Generate It?
Learn what LLMs.txt is, its SEO importance, and how it differs from robots.txt. Get a step-by-step guide to generate it and optimize your site for AI-driven search success.
This llms.txt generator turns your sitemap into a markdown index of up to 5,000 pages, each with an AI-written summary. llms.txt is not a ranking signal — Google confirmed in June 2026 it has no effect on Search or AI Mode — but it gives an agent pointed at your site text instead of a navigation menu. Free, no signup.
The generator is sitemap-driven, not a blind crawl. It reads robots.txt for a sitemap declaration, falls back to the common paths when there isn't one, expands index sitemaps into their child URLs, then fetches each page and passes the extracted text to a language model that writes one title and one description per URL. Every URL resolves in front of you, and failures are named rather than quietly dropped — so a thin result tells you something real about your sitemap.
A bare domain is enough — example.com. You don't need to own the site, verify anything, install a snippet, or sign in.
robots.txt is checked first, then the common sitemap paths, then any index sitemap is expanded into its children. Up to 5,000 pages per domain are queued and processed in front of you.
Each page returns an AI-written title and description in llms.txt format. Fix the summaries that read wrong, then copy the file or download it and upload it to your site's root as /llms.txt.
Everything in the generator is free and needs no sign-up. The marked items are the follow-on checks a free SEOmator account unlocks — the ones that measure whether AI systems actually reach your site, which the file itself cannot tell you.
The tool finds your sitemap from robots.txt or the common locations and expands index sitemaps into their children. Nothing to paste, nothing to configure.
Large documentation sets and content archives are processed whole, not truncated to a sample you then have to finish by hand.
Each page gets its own title and one-line description generated from the page's actual text — the part that makes an llms.txt useful instead of a second sitemap in markdown.
Output follows the llms.txt structure the proposal defines — an H1 site name, a blockquote summary, then H2 sections of described links. No reformatting before you ship it.
Watch each page resolve as it's processed, with failures named individually and a cancel control if a run turns up more than you wanted.
The generated file is editable before you take it — trim pages, rewrite summaries — then download it or copy it straight to the clipboard.
Run all 251 audit rules against the domain, including the AI/GEO readiness checks that look at what agents can actually read on your pages.
Agent Analytics reads your server logs and names the bots hitting your site, so you can check whether anything ever requests the file you just shipped.
Monitor whether your brand shows up in AI answers over time — the outcome llms.txt is usually bought for, measured directly instead of assumed.
Create a free SEOmator account to audit the site, watch the crawlers that reach it, and track where it gets cited.
Sign up freeI'll be blunt about this file, because the pages selling you a generator usually aren't. llms.txt is not a ranking signal. Google said as much in June 2026, and no major AI engine has confirmed that it reads one at retrieval time. If you are shipping it to get cited in ChatGPT or AI Overviews, you will be disappointed — and the wider web is making a matching mistake in the other direction: across the 109,440 top domains we scanned on 20 July 2026, 81.2% write AI-crawler rules into robots.txt while only 6.6% serve agents anything readable. Everyone is writing rules for agents nobody is ready to serve. Where I do use llms.txt is narrower and real: when a person or an agent is handed your URL, /llms.txt is the one address that answers with your site as text instead of a navigation menu. That's worth twenty minutes. It isn't worth a strategy, and anyone telling you otherwise is selling the file, not using it.
The only honest measure of this file is whether anything requests it. Across the 109,440 domains we scanned on 20 July 2026, 81.2% had written AI-crawler rules but only 6.6% served agents readable markdown — rules without readiness. Ship the file, then read your access logs or Agent Analytics for hits on /llms.txt. If nothing fetches it in a month, you have your answer and you've lost twenty minutes, not a quarter.
More than one in three crawler requests on the sites we monitor was rejected outright by June 2026 — a 36.4% 4xx rate. A WAF configured to block unfamiliar user agents blocks the agents you wrote the file for. Fetch your own /llms.txt with a non-browser user agent before you assume it's live.
Access decisions outrank index files. Across 4,257 robots.txt files we parsed on 20 July 2026, GPTBot (773 files) and ClaudeBot (629) were the most-named agents on the web — allow or deny is a real, enforced signal in a way llms.txt is not. Check yours with the robots.txt tester first.
The generator will happily queue every URL in your sitemap. That's the starting point, not the file. Delete tag archives, paginated listings and anything you wouldn't want quoted back at you — an index that points at your weakest pages is worse than no index.
The generated descriptions come from each page's own text, which is accurate and a little flat. The twenty pages you'd actually want an assistant to describe deserve a human sentence each. That's where the file stops being a sitemap in markdown.
Serving clean, readable content matters far more than the index pointing at it. Run a GEO audit to see what agents can parse on the pages themselves, and check whether you're being cited in AI answers at all — that's the number llms.txt is usually bought to move, and the only one worth watching.
An llms.txt file is a markdown document at the root of a website — yoursite.com/llms.txt — that summarizes what the site is and links to its most useful pages, each with a short description. Jeremy Howard of Answer.AI proposed the format in September 2024 to solve a specific problem: a language model handed a normal web page gets navigation, cookie banners, ads and script tags along with the content, and has a limited context window to spend on the mess.
The structure is deliberately small:
That's the whole specification. An optional llms-full.txt carries the expanded content of those same pages for cases where a model should read the material rather than navigate to it.
It is worth being precise about what the file is not. It does not control access — that is robots.txt, and it is enforced. It does not drive indexing — that is sitemap.xml. It is a summary written for a reader that prefers text, published at a predictable address.
Not for the reason most people ship it. Two years after the proposal, the evidence is one-directional and worth stating plainly, because almost every page selling a generator declines to.
Google does not use it. Google updated its AI-optimization guidance in June 2026 to state that llms.txt has no effect — positive or negative — on Search rankings or AI Mode. John Mueller has described the standard as speculative more than once, and has pointed at WebMCP as the direction Google finds more promising.
Almost nothing fetches it. Ahrefs analyzed 137,000 sites and found roughly 97% of published llms.txt files are never requested by AI crawlers at all. A file nothing fetches cannot influence anything downstream of the fetch.
No provider has confirmed reading it. Plenty of well-known companies publish one — Anthropic, Cursor and Mux among them. Publishing is visible and cheap. What no major AI provider has documented is that its models or crawlers retrieve and use the file when answering a question. Those are different claims, and coverage of llms.txt routinely blurs them.
So the honest summary: llms.txt is a real, published format with real adopters on the publishing side and unconfirmed support on the consuming side. If your goal is ranking or citation, this file is not the lever.
There is one use that survives all of the above, and it is the reason our own generator exists.
When a person or an agent is handed your URL — a developer pasting your docs link into a coding assistant, a researcher pointing a model at your site, an agent following a link a user gave it — /llms.txt is a predictable address that answers with your site as text. No rendering, no navigation, no cookie interstitial. That is a small, concrete benefit, available today, that does not depend on anyone adopting a standard.
It shows up most for documentation, API references and developer tools, which is exactly the profile of the companies that publish one. If you run a docs site, this is twenty minutes well spent. If you run a local services site, the same twenty minutes in your page titles will do more.
This is where our own data says something the coverage doesn't. We ran the SEOmator audit engine's AI-readiness checks across 109,440 top domains on 20 July 2026:
| Signal | Domains | Share |
|---|---|---|
| robots.txt present | 85,602 | 78.2% |
| robots.txt AI rules | 88,838 | 81.2% |
| sitemap.xml | 77,079 | 70.4% |
| Markdown for agents | 7,268 | 6.6% |
| Content-usage signals | 6,910 | 6.3% |
| MCP server card | 285 | 0.26% |
Source: SEOmator crawling data, scan of 2026-07-20 (109,440 top domains).
Four out of five sites have written rules telling AI agents what they may and may not do. Fewer than one in fifteen serve those agents anything readable when they comply. The web is writing rules for AI agents it isn't ready to serve.
That gap is the honest case for llms.txt, and it is a narrower case than the marketing around the file suggests. Publishing one moves you into the 6.6% for the cost of an afternoon. It does not, on its own, make you citable.
The same scan's robots.txt parsing is worth reading alongside it. Across 4,257 robots.txt files parsed on the same date, GPTBot appeared in 773 and ClaudeBot in 629 — the two most-named agents on the web. AI opt-out is now the single most common reason a site writes crawler rules at all. Those rules are enforced by every major crawler. llms.txt is not enforced by anyone. Spend your attention accordingly, and check what yours currently says with the robots.txt tester.
The generator at the top of this page is sitemap-driven rather than a blind crawl, which matters for interpreting what comes back.
example.com. You don't need to own the site or verify anything.yoursite.com/llms.txt, served as plain text.A thin or empty result is diagnostic, not a bug: it usually means the domain has no discoverable sitemap. That is worth fixing regardless of what you think of llms.txt, since it affects the crawlers that demonstrably do matter. The sitemap finder will tell you what's discoverable.
# Acme Docs
> API and SDK documentation for Acme, a payments platform for marketplaces.
## Getting started
- [Quickstart](https://acme.dev/docs/quickstart): create an account, get a key, make a first charge in about ten minutes.
- [Authentication](https://acme.dev/docs/auth): key types, rotation, and the difference between test and live mode.
## API reference
- [Payments API](https://acme.dev/docs/api/payments): create, capture, refund and dispute endpoints with request and response shapes.
- [Webhooks](https://acme.dev/docs/api/webhooks): event types, retry behaviour, and signature verification.
## Optional
- [Changelog](https://acme.dev/changelog): dated list of API changes and deprecations.
Note what is missing: no tag archives, no paginated listings, no marketing pages, no "Solutions" hub. Roughly a dozen entries, each described in a sentence that says what the reader will find.
Four patterns account for most of the bad files we see.
Shipping the whole sitemap. The generator will queue every URL it finds. That's the raw material, not the file. An index that points at 2,000 pages including your tag archives is a sitemap with different punctuation — it gives a model no signal about what matters. Cut it to the pages you'd genuinely want quoted.
Leaving the generated summaries untouched. The descriptions are written from each page's own text, which makes them accurate and a little flat. For your top twenty pages, write the line yourself. That's where the file stops being mechanical.
Not checking that it's reachable. More than one in three crawler requests on the sites we monitor was rejected outright by June 2026 — a 36.4% 4xx rate across our panel. A WAF or bot-management rule configured to block unfamiliar user agents will block the agents you wrote the file for. Fetch your own /llms.txt with a non-browser user agent before assuming it is live.
Letting it rot. A stale index pointing at 404s is worse than no index: every dead link in it wastes the one request you were hoping for. Re-generate when the structure changes — a new docs section, a renamed product, a batch of retired pages. Quarterly is a reasonable default for a stable site.
If the actual goal is showing up in AI answers, llms.txt is a footnote. Three things matter more, in order.
Access. Decide per bot purpose rather than with a blanket rule. Training crawlers, retrieval bots and user-triggered agents are three different things: blocking the retrieval and user-triggered agents removes you from answers where you would otherwise be cited, which is usually not what a site means to do. Those decisions live in robots.txt, and they are enforced.
Readability. A page an agent cannot render is a page it cannot quote. Content that only exists after JavaScript executes is invisible to a large share of the bots fetching it. This is where the 6.6% figure above bites: sites write the rules and then serve agents markup they can't use. A GEO audit checks what agents can actually parse on your pages.
Measurement. Everything above is a hypothesis until you look. Server logs — or Agent Analytics — tell you which bots reach you and what they fetch, including whether anything ever requests the file you just published. An AI visibility check tells you whether your brand appears in AI answers at all. That last number is the one llms.txt is usually bought to move, and the only one worth watching.
Generate the file. It costs twenty minutes and it puts you in a small minority of sites that serve agents something readable. Then go and do the work that the evidence actually supports.
This is the audience where the file genuinely lands. Developers paste doc URLs into coding assistants all day; a clean /llms.txt means the assistant gets your API reference as text instead of a rendered sidebar. Anthropic, Cursor and Mux all publish one for exactly this reason.
A client read that llms.txt is the next big thing and wants it. Generate one in five minutes, tell them honestly what it does and doesn't do, and spend the rest of the retainer on the crawl and content work that moves AI visibility.
You want the cheap experiment on the record before anyone asks why it wasn't tried. Ship the file, log whether anything fetches it, and you have evidence rather than an opinion the next time it comes up in a planning meeting.
No account, no card, and no "we processed 400 pages — sign up to see them" step. The file you generate is the whole file.
Most free generators sample the first handful of URLs they find. A documentation set or a content archive comes back whole here, which is the difference between a draft and a file you can ship.
Each entry carries a description generated from that page's own text. An index of bare URLs is just your sitemap with different punctuation.
Trim the pages you don't want quoted and rewrite the lines that matter, in the tool, before you download. Curation is most of the value here.
This page tells you llms.txt is not a ranking signal, which costs us a conversion and saves you a bad assumption. The GEO audit is where the work that does move AI visibility starts.
One of 40 free SEO tools, from a robots.txt tester to a complete site audit.
llms.txt is a proposed standard for a markdown file at the root of a site — yoursite.com/llms.txt — that summarizes the site and links to its most useful pages with a short description for each. The format is an H1 site name, a blockquote summary, then H2 sections of described links. It was proposed by Jeremy Howard of Answer.AI in September 2024 as a way to hand language models a clean text version of a site instead of rendered HTML full of navigation, ads and scripts.
It's a real, published proposal with real adopters on the publishing side — Anthropic, Cursor and Mux all serve one — but it is not an adopted standard on the consuming side. No major AI company has confirmed that its systems read llms.txt when answering a question, and Google has stated plainly that it does not use the file. So: real file, real format, genuinely useful in the narrow case below, and not the industry standard that some coverage implies.
No. Google updated its AI-optimization guidance in June 2026 to state that llms.txt has no effect — positive or negative — on Search rankings or AI Mode, and John Mueller has repeatedly described the standard as speculative and unused by Google. Independent analysis points the same way: Ahrefs' study of 137,000 sites found roughly 97% of llms.txt files are never fetched by AI crawlers at all. If your goal is ranking or citation, the file is not the lever.
For most sites, it's a twenty-minute job with a small, specific payoff — worth doing, not worth planning around. The payoff is real in one case: if people or agents are handed your URLs directly, /llms.txt is a predictable address that returns your site as text. Documentation, API references and developer tools benefit most. If you run a local business site or a small blog, the same twenty minutes spent on page titles or internal links will do more.
Keep it short and curated. An H1 with your site or product name; a blockquote of one or two sentences saying what you do; then H2 sections grouping your genuinely useful pages — documentation, guides, pricing, API — each as a markdown link with a short description of what the page covers. Leave out tag archives, paginated listings, thin pages and anything you would not want quoted back at you. An optional llms-full.txt can carry the expanded content for the same pages.
Use the generator at the top of this page. Enter your domain, and it finds your sitemap from robots.txt or the common paths, processes up to 5,000 pages, and writes a title and description for each one. Edit the result — trim the pages you don't want listed, rewrite the summaries that matter — then download or copy the file and upload it to your site's root so it resolves at yoursite.com/llms.txt.
Yes, and you should. The generated file is a starting point, not a finished one — it will list whatever your sitemap contains, at whatever quality the page text supports. Edit it in the tool before downloading: cut the URLs that don't earn a place, and write human descriptions for your most important twenty pages. A curated 40-link file is more useful to a model than an exhaustive 2,000-link one.
They solve different problems and only one of them is enforced. robots.txt controls crawler access and is honored by every major crawler — it is the file that actually decides whether AI systems can read you. sitemap.xml lists your URLs for indexing and is read by search engines. llms.txt proposes a human-readable markdown summary for language models, and support for it is voluntary and largely unconfirmed. Fix robots.txt first; llms.txt is the optional extra.
It depends on what you're optimizing for, and it's a bigger decision than llms.txt. Blocking training crawlers protects your content from being absorbed; it also removes you from the corpus that answers get built from. Blocking retrieval and user-triggered agents — the ones fetching a page because someone asked about you — removes you from answers where you'd otherwise be cited. Across 4,257 robots.txt files we parsed on 20 July 2026, GPTBot and ClaudeBot were the two most-named agents on the web, so most sites are making this call deliberately. Decide per bot purpose, not with one blanket rule.
No, and it's worth separating publishing from consuming. Plenty of well-known companies publish an llms.txt — that part is easy and visible. What no major provider has confirmed is that its models or crawlers fetch and use the file when answering a question. Treat every claim that a specific assistant "reads llms.txt" with suspicion unless that provider has documented it.
Whenever the structure it describes changes — a new documentation section, a renamed product, a batch of retired pages. Quarterly is a sensible default for a stable site. A stale index that points at 404s is worse than no index: if something does fetch the file, every dead link in it is a wasted request.
The convention is one file at the root. It can link out to additional markdown files for specific sections, which keeps the main file short while still exposing depth — that's what llms-full.txt is for. Subdirectory files like /docs/llms.txt have no defined discovery mechanism, so nothing will find them unless you link to them.
It works better on a large site, but only if you curate. The generator handles up to 5,000 pages per domain, which covers most documentation sets and archives. Shipping all 5,000 is the mistake: an index is only useful when it points at the pages worth reading. Generate the full list, then cut it down to the sections you'd actually want summarized.
Be short, be specific, and describe every link — a bare URL tells a model nothing your sitemap didn't. Write in plain language rather than marketing copy, since the file is read as text and not rendered. Verify the file resolves as plain markdown at the root, check that nothing in front of your site blocks non-browser requests to it, and test it by pasting the contents into an assistant to see whether it can answer questions about your site from that alone.
No. The generator on this page produces the file from your domain, and publishing it means uploading one text file to your site's root — the same step as adding a favicon or a verification file. Most CMS and hosting platforms let you do that from an admin panel or a file manager without touching code.
You can, and it costs little, but be realistic about the return. The file earns its keep where people hand URLs to assistants — documentation, APIs, developer tools. A blog or a store can list its main categories and best guides the same way, but for those sites the effort is better spent on things AI systems demonstrably do read: clean, crawlable pages and clear on-page structure.
One concrete benefit and two conditional ones. Concretely: a predictable text address that returns your site as markdown, useful whenever a person or an agent is pointed at your URL. Conditionally: if support ever becomes real, sites that already publish a maintained file start ahead; and the act of writing one forces a useful audit of which pages you would actually put in front of a model. What it is not is a ranking or citation lever today.
Work on what the systems verifiably read. Make sure retrieval and user-triggered agents aren't blocked in robots.txt, since access decisions are enforced in a way llms.txt is not. Make your pages readable without JavaScript, because a page an agent can't render is a page it can't quote. Then measure: run a GEO audit to see what agents can parse on your pages, and track whether your brand shows up in AI answers at all. That last number is the one llms.txt is usually bought to move.
Deeper reading from the SEOmator blog on what this tool measures.

Learn what LLMs.txt is, its SEO importance, and how it differs from robots.txt. Get a step-by-step guide to generate it and optimize your site for AI-driven search success.

Only 46% of AI-crawler requests now get a 200 OK, more than a quarter are blocked or throttled, and which bot dominates flips completely by country.

Mistral crawls 3,389 pages for every referral it sends back — now the most extractive AI bot on the web, ahead of Anthropic (2,237:1) and OpenAI (217:1). I cross-checked Cloudflare Radar's global network against SEOmator's own 500+ site panel to rebuild the crawl-to-refer ratio for July 2026.
The file is the easy part. Run a free GEO audit to find what's blocking you.