llms.txt is a Markdown file at the root of your domain that gives AI models a curated summary of your site: who you are, what you do, and which pages matter. It is not a replacement for robots.txt and does not guarantee citations, but it is a low-cost, low-risk signal that takes about an hour to implement correctly.
What llms.txt actually is
llms.txt is a Markdown file placed at the root of a domain, at /llms.txt, that gives large language models a curated summary of what a site is and which pages carry its most useful information. It was proposed in 2024 as a way to help AI systems read a site efficiently, rather than crawling and guessing.
The reasoning is straightforward. A modern web page is mostly navigation, scripts, styling and boilerplate. A language model working from raw HTML spends most of its context on markup that carries no meaning. llms.txt hands over the substance directly: who you are, what you do, and where the real content lives.
What llms.txt is not
It is not a replacement for robots.txt, and it is not an access control mechanism. robots.txt tells crawlers what they may fetch. llms.txt tells models what is worth reading. If you want to block an AI crawler you do that in robots.txt, with a user-agent rule for GPTBot, ClaudeBot, PerplexityBot or whichever agent you mean.
It is also not an official standard adopted by every AI provider. Support varies, and no major engine has committed to reading it as a ranking input. Treat it as a low-cost, low-risk signal that costs an hour to implement, not as a guaranteed route to citations.
What belongs in the file
The format is deliberately simple. An H1 with the site or organisation name, a blockquote giving a one-paragraph summary, then H2 sections containing lists of links with short descriptions. Everything is Markdown, and everything should be written for a reader who knows nothing about you.
The four things worth including
- Identity. The organisation name, what it does, where it operates, and who leads it. This is the part models use to resolve your brand as an entity.
- Key facts. Founding details, legal entity, location, contact address, and the services or products you actually offer. Keep these identical to what your structured data says.
- Primary pages. A curated list, not a sitemap. Ten to twenty pages that genuinely answer questions, each with a description of what the reader will find there.
- Nothing you would not publish. The file is public and will be fetched. It is not a place for pricing you do not advertise or internal notes.
A working example
This is the shape of a well-formed file. Keep the summary tight, make every description earn its place, and never list a URL that redirects or is set to noindex.
# EXAMPLE COMPANY
> Example Company is a Colombo-based accountancy practice specialising in
> audit and tax compliance for small and medium Sri Lankan businesses.
## Key facts
- Founded: 2015
- Legal entity: Example Company (Pvt) Ltd
- Location: Colombo 03, Sri Lanka
- Services: Statutory audit, tax filing, payroll, company secretarial
- Contact: hello@example.lk
## Primary pages
- [Services](https://example.lk/services): Full list of accountancy services
- [Tax filing guide](https://example.lk/tax-guide): How SME tax filing works in Sri Lanka
- [Contact](https://example.lk/contact): Book an initial consultation
The mistake almost everyone makes
The single most common error is letting the file go stale. A page gets retired, redirected or set to noindex, and the llms.txt entry still points at it. You are then actively directing the exact crawlers you want to impress toward a dead URL, which is worse than having no file at all.
Treat llms.txt as part of your deployment checklist. Whenever a page is removed or its URL changes, the file gets updated in the same commit. It takes seconds and prevents the one failure mode that makes the file counterproductive.
How to check it is working
- Fetch
https://yourdomain.com/llms.txtin a browser and confirm it returns plain text, not a 404 or an HTML error page. - Check the server sends it as
text/plain. Some hosts serve unknown extensions as a download, which some fetchers reject. - Click every URL in the file and confirm each returns a 200 and is indexable.
- Ask an AI assistant a question your site should answer, and see whether it surfaces you. This is not a controlled test, but repeated over weeks it shows direction.
None of this guarantees a citation. What it does is remove the excuses: if a model does not cite you, it will not be because your site was unreadable.