
What Is llms.txt? A Practical Spec for Site Owners
When people search for llms.txt, they usually want a straight answer: is this a new crawl-control file, an SEO trick, or something else entirely? It is a proposed convention—a short Markdown document you publish so language-model systems can read a human-authored brief about your site instead of guessing from noisy HTML.
The file is not a robots.txt clone. Robots.txt still governs whether a crawler may fetch URLs. llms.txt is closer to a crawl contract for answer engines: who you are, which docs are canonical, what to skip, and how to cite or summarize you. It also is not a citation generator; it does not mint links or force models to quote you. Adoption is uneven, parsers differ, and nothing here overrides law, robots rules, or authentication.
This article is an operator spec. You will get placement, typical fields, how it collides with existing crawl signals, who should own the file, and how it relates to answer-engine optimization—without treating a draft convention as a ranking guarantee.
- llms.txt is optional Markdown at a well-known path that briefs AI systems on your site, not a replacement for robots.txt.
- It describes purpose, key pages, and usage notes; it does not grant crawl rights or auto-create citations.
- Place it where fetchers expect it, keep it truthful, and resolve conflicts with robots and sitemaps in robots’ favor.
- Treat it as content ops plus crawl policy, not as a magic AI-visibility plugin.
- Generators can draft structure; a human still owns accuracy and legal claims.
llms.txt Is a Curated Map, Not a Crawl Policy
That distinction is the whole point of the file. robots.txt is a crawl policy: it tells automated agents which paths they may fetch and which they must skip. sitemap.xml is an inventory: it lists URLs you want discovered, often with last-modified hints. Meta robots and X-Robots-Tag control indexation in search. llms.txt sits beside those tools. It does not allow or deny crawling, it does not replace your sitemap, and it does not flip a ranking switch. It is an optional, human-readable Markdown document at a well-known URL that you can use to say what the site is, which pages you consider canonical or durable, and how you would like answer engines to treat that material.
Think of it as a curated map for LLM-oriented crawlers and answer engines—not a contract they are bound to honor. You nominate context: product names, documentation hubs, policy pages, the URLs that still make sense if someone asks a question six months from now. You do not command citations. Treating “how to get cited by ChatGPT” as the job description for this file will only disappoint you. Citation is an outcome of retrieval, grounding, and the model’s own policies. A well-written llms.txt can reduce guesswork about which pages matter; it cannot force an answer engine to quote you, rank you, or even fetch you.
That is why it is useful as an AEO control surface and dangerous as a legal stand-in. Answer engine optimization here means you choose the durable URLs and the surrounding context you want systems to prefer when they do look. Ignoring the file is allowed. Some crawlers will never request it; others will read it and still pick different pages. If you treat llms.txt like a robots rule or a terms-of-service clause, you invent false confidence: you assume a permission model that does not exist. Keep robots.txt for access, sitemaps for inventory, legal pages for rights, and llms.txt for a voluntary map—then you will use it for what it actually is.
Map, not mandate — llms.txt nominates what your site is and which URLs should matter to answer engines; it does not allow, deny, index, or legally bind anyone, and it cannot make ChatGPT cite you.
What a Valid llms.txt File Actually Looks Like
That map only works if the file itself is easy for a model to parse, which is why the emerging convention is so specific. Think of llms.txt as an operator contract written in plain Markdown: a human-readable briefing, not a list of crawl rules. A valid file starts with the site name as a single H1, followed by a short blockquote that states what the site is and who it is for. After that come H2 sections of links, each URL paired with one line of description so an answer engine can tell why that page exists without fetching it first.
Keep the primary list tight
Nominate the durable pages that should represent you: product overviews, docs hubs, pricing, policies, and evergreen explainers. Optional secondary resources—changelogs, older posts, niche references—belong in their own later H2 so the first list stays short. If everything is “important,” nothing is.
- Plain Markdown only, saved as UTF-8.
- No HTML, no scripts, no embeds, no tracking parameters on listed URLs.
- Canonical, stable links—not session IDs, UTM tags, or faceted duplicates.
- Keep the whole file small enough to ingest in one pass: a few dozen well-described links beat hundreds of unexplained ones.
Density matters as much as size. One-line descriptions should name the page’s job (“How billing works for annual plans”), not repeat the title. If a model cannot finish the file, your curation never arrives.
If you paste Allow and Disallow lines into Markdown and call it llms.txt, you have not written the contract. You have restated crawl policy in the wrong file. The next practical question is where this file lives on the origin—and how it sits beside robots.txt without pretending to replace it.
The contract — A valid llms.txt is UTF-8 Markdown with an H1, a short summary, and sparse H2 link lists—never HTML, never tracking URLs, and never a robots directive dump.
Where llms.txt Lives—and Why robots.txt Can Quietly Kill It
That Markdown contract only works if fetchers can actually retrieve it. Ship the file at the site root as https://example.com/llms.txt (swap in your real host). It should return a stable 200, sit behind no login or paywall, and advertise a text/plain or Markdown content-type. Treat that URL as a well-known address, not a marketing landing page: no query strings, no HTML wrapper, no tracking parameters.
Fetch failures are usually configuration, not the spec. A 301 or 302 chain, a www versus apex split, or a CDN that serves a stale 404 can hide the file from answer-engine crawlers that only try the canonical root. Cache headers that lock an empty or HTML error body in place have the same effect. Confirm the live response yourself: one hop, 200, UTF-8 Markdown, identical on every hostname you care about.
The robots.txt contradiction
llms.txt does not override robots.txt. If you Disallow the user-agents those engines use, they may never fetch the map you published for them. Expecting AI visibility from a curated file while blocking the same crawlers is a policy collision, not a ranking trick. The same silent failure happens when URLs you list are noindex, blocked by a robots meta tag, or gated behind login: the map points somewhere the engine is not allowed to read. Fix the crawl and indexation rules first, then nominate pages.
Some sites also publish an optional llms-full.txt: a longer companion with extra context. It is not a substitute for the small, curated root file. Keep the well-known /llms.txt short and ingestible; treat the full variant as overflow, not the primary map.
Reachability first — A valid Markdown map at the wrong URL, behind a block, or pointing at noindex pages never reaches the engines you wrote it for.
What belongs in llms.txt (and what generators dump in by mistake)
Once the file is reachable, the next mistake is stuffing it. A generator that walks every published URL will happily list tag pages, page 2 of a blog, and login-gated dashboards. That is an inventory dump, not a map. The spec’s job is nomination: a short list of durable, high-intent canonicals you actually want an answer engine to treat as the site’s spine.
Start with cornerstone guides, product and docs hubs, and original research that still reads true a year later. Skip the daily recap, the thin comparison post, and anything that exists only to catch a long-tail keyword. If a URL would embarrass you as a citation next to your brand name, it does not belong in the file.
Write descriptions as claims, not slogans
Each link needs one factual line a retriever can use: what the page is, who it is for, and what it covers. “Pricing for the Pro plan, including seats and overage” beats “Unlock growth with our revolutionary platform.” Marketing copy is noise in this format. Concrete nouns, versions, and scope are signal.
Leave the crawl junk out—especially on WordPress
Exclude login walls, thin tag indexes, author archives, faceted search, and paginated lists. WordPress sites inflate fastest: /category/, /tag/, /page/2/, and date archives look like content to a crawler and like clutter to a model. Those URLs duplicate or fragment the same posts; nominating them trains the engine on the wrong surface.
Omission is a spec decision, not a gap. A short file that matches what you want cited beats a complete dump of the crawl. If a page is missing, you chose not to endorse it—same as leaving a product off a sales sheet. Generators that “complete” the list for you undo that choice.
Curate, don’t dump — Nominate a handful of canonical, citable pages with factual one-liners; a small curated map is the product, not a sitemap reprint.
Who owns llms.txt after you ship it
Once the map is live, the next failure mode is not a bad URL list—it is nobody owning the file. If “someone’s plugin” is the owner, the file dies the first time a theme update, CDN rule, or sitemap export overwrites it. Assign a named person or team (search, docs, or content ops) who can approve additions, reject dumps, and keep the Markdown honest.
Refresh on a meaningful publish or information-architecture change: a new canonical guide, a retired product line, a docs hub move. Do not run a noisy daily regenerate that re-scrapes tags and pagination. The whole point of a curated map is that it stays small and intentional; automation that rewrites it every night undoes that.
Static file versus CMS-managed file
A committed static file in version control is usually safer: you see every line in review, deploys are explicit, and a plugin cannot silently append archives. A WordPress-managed file can be safer when non-engineers must update links without a deploy, but only if generation is gated—templates that dump every post type will recreate the crawl dump you already rejected. Prefer static when the site is mostly durable docs; prefer a gated CMS flow when editors own the canonicals and engineering will not touch Markdown.
Treat origins as separate maps. Staging should not advertise production URLs, and production should never serve a staging copy. Each public hostname—apex versus www, brand microsite, locale subdomain—needs its own well-known file if answer engines will fetch that origin. Do not assume one file on the marketing domain covers docs.example.com or a regional host.
That gate is the operational counterpart of omission: if a URL was not nominated, it does not belong. Generators can draft candidates; they should not write the live file unattended.
Ownership — llms.txt stays useful only while a named owner treats it as a curated, per-origin artifact—updated on real IA change, never auto-scraped into existence.
How to tell whether anything actually read your llms.txt
Once a named owner has shipped a curated map, the next question is operational, not ceremonial: did any machine fetch it, and could it even use what it found? Re-request the live file and walk the nominated URLs as they exist today—not as they were on publish day. Drift, host splits, and pages that now redirect into duplicates will poison logs and retrieval alike; fix the map before you interpret silence.
Proof of interest lives in server or CDN logs as GET requests for /llms.txt, not in a screenshot of a chatbot citing your brand. User-agent strings vary and are easy to spoof, so treat a fetch as “something asked for the file,” not as “ChatGPT ranked you.” Day-one silence is normal; many answer engines cache, sample, or skip the convention entirely.
Failure modes that look like “nobody read it”
- AI crawlers Disallowed in robots.txt, or the file hosted on the wrong host (www vs apex, locale vs brand).
- A stale CDN copy serving an old map, or a 301/302 that never settles on the well-known path.
- Nominated URLs that are noindexed, behind auth, or stuffed into a file too long to ingest usefully.
A fetch still is not a ranking or citation event. Use log hits only as a reachability check, then judge the map by whether the nominated pages already state what you want said. Silence can mean the engine skipped the convention; the durable work is still crawlable source pages plus a short, honest file.
Health check — Verify the file and its URLs first, then watch logs for fetches—and remember that ignore is allowed, so the real product is still crawlable pages plus a short, honest map.
Key Takeaways
Put a short, owned llms.txt at your site root, pair it with crawlable canonical pages, and check server logs to see whether answer engines actually fetch it.