How does llms.txt actually work?
The format is deliberately simple. It's a markdown file with:
- An H1 with your site or project name
- A blockquote summarizing what you do
- H2 sections grouping links (e.g., "Services," "Guides," "Pricing")
- Each link with a one-line description of what's on that page
An optional companion, llms-full.txt, contains the full text of your key pages in one file, so a model can ingest everything in a single request. There's no registration and no gatekeeper — you just publish the file, the same way robots.txt works. Because it's voluntary and unenforced, nothing on your site changes when you add it; you're simply offering AI systems a cleaner entry point and hoping the ones that matter choose to use it.
The theory: LLMs have limited context windows and struggle with cluttered HTML. A clean, curated index lets them grab your canonical answers — your pricing, your service definitions, your FAQs — accurately instead of hallucinating or citing a competitor's description of you. When it works, the file is the difference between a model quoting your own words about your business and a model guessing.
Does llms.txt actually work? What the picture says
Here's the honest read as of 2026, and it cuts both ways.
The adoption side is real. Over the past year, major documentation platforms, SaaS companies, and dev tools have shipped llms.txt across large swaths of the web. Mintlify, for example, rolled the file out across the documentation sites it hosts, publishing it by default rather than site by site. If you spend time in developer docs, you now trip over llms.txt constantly — the format won the "should we publish one?" debate on momentum alone.
The usage side is sobering. Today's major AI crawlers largely ignore the file. Google's John Mueller has publicly compared it to the old keywords meta tag, noting that "none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it)." His deeper objections are worth understanding: a model that has already downloaded your real content has no reason to trust a separate file describing it, and anyone could show one set of pages in llms.txt while serving something else to users — cloaking for LLMs. Neither Google nor OpenAI has formally committed to consuming the file.
So why publish one at all? Three reasons we still ship it inside our generative engine optimization engagements:
- The files that do get crawled share traits — big documentation sets, frequently updated content, and sites that AI agents already visit. If agents visit you, the file has a consumer.
- It costs an hour. The asymmetry is absurd: near-zero cost, a real (if minority) chance of payoff, and no downside risk.
- Agent traffic is the trend line. AI search referrals and agentic browsing — AI tools acting on behalf of users — are climbing fast, and files built for machine readers are a bet on where traffic is going, not where it's been. The cost of being early is one hour; the cost of being late could be a category.
— "llms.txt is a lottery ticket that costs an hour. The mistake isn't buying the ticket — it's believing the ticket is the strategy." — Thomas, Founder of AISEO USA
How to implement llms.txt properly (5 steps)
In the audits we run, most llms.txt files fail the same way: they're auto-generated sitemaps renamed, with no curation and no descriptions. That's how a file ends up ignored. Here's the implementation that gives the file a job to do:
- Curate ruthlessly. List 10–40 pages that answer buying and usage questions — services, pricing, comparisons, FAQs. Not every blog post. An llms.txt that mirrors your sitemap tells a model nothing your sitemap didn't.
- Write real descriptions. Each link gets one sentence stating what question the page answers. This is the metadata a model uses to decide relevance.
- Make the destination pages extraction-ready. The file points; the pages persuade. Answer-first openings, question H2s, and clean markup matter more than the index itself — that's core LLM SEO, not file formatting.
- Ship llms-full.txt if you have docs or deep guides. Full-text ingestion is where documentation-heavy sites see actual agent requests.
- Check your server logs monthly. Grep for GPTBot, ClaudeBot, PerplexityBot, and friends hitting the file. If they're reading it, expand it. If not, you've lost nothing — and you've confirmed where to spend instead.
Freshness applies here too. A stale llms.txt pointing at stale pages is doubly invisible: AI engines lean toward recently updated sources, and a file that hasn't been touched in a year signals a site that hasn't either. Treat the index and its destination pages as one living asset — update them together, or don't bother.
What matters more than llms.txt?
This is the section that saves you money. If your goal is showing up in ChatGPT, Perplexity, and Google's AI answers, llms.txt is a tiny slice of the job. The rest, in rough priority order:
| Priority | Lever | Why it outranks llms.txt |
|---|---|---|
| 1 | Answer-first content structure | It's what engines actually extract and cite |
| 2 | Entity consistency across the web | Models recommend brands they can verify from multiple sources |
| 3 | Third-party mentions and lists | Reference and community domains dominate AI citations (Semrush) — mentions and lists get you into them |
| 4 | Schema markup (FAQ, Service, Organization) | A machine-readable layer engines demonstrably parse |
| 5 | Content freshness | AI engines favor recently updated pages when choosing what to cite |
| 6 | llms.txt / llms-full.txt | Cheap, unproven, worth an hour |
Notice what this list is: it's AI search optimization as a system. One data point makes the priorities concrete — Ahrefs found that 62% of AI Overview citations come from pages ranking outside the top 10 organic results. In other words, being cited by AI is not the same game as ranking, which is exactly why entity strength, mentions, and clean extractable content beat a text file you host yourself.
So, does llms.txt work? Only as well as the system around it. Publishing a text file doesn't make a business citable any more than printing a menu makes food good. But when the system exists, the file gives well-behaved agents a clean front door.
We'd rather tell you that plainly than sell you a "proprietary llms.txt deployment" — and while we're being plain: no one can guarantee AI citations either. Anyone who does is lying. Our work raises the probability, and we show you the receipts monthly.
What does a good llms.txt actually look like?
Here's a stripped-down example for a service business, so you can see how little ceremony the format needs:
# AISEO USA
> AI SEO agency helping US businesses rank on Google and get cited
> by ChatGPT, Perplexity, Gemini, and AI Overviews.
## Services
- [AI SEO Services](https://aiseousa.com/ai-seo-services): What our AI SEO programs include and who they fit
- [SEO Packages & Pricing](https://aiseousa.com/seo-packages-pricing): 2026 pricing guide — market rates and how we quote
## Guides
- [How to Rank in ChatGPT](https://aiseousa.com/blog/how-to-rank-in-chatgpt-2026-playbook): The 7-step 2026 playbook
Three details worth copying: the blockquote states what the business is in one breath (that's the line a model reuses when describing you), every link answers a distinct question, and there's nothing here a human wouldn't also find useful. If your llms.txt would embarrass you as a public "start here" page, it's not curated enough.
Should your business publish one?
Yes, with correct expectations. Publish it if any of these apply: you have documentation or deep guides, you sell something researched via ChatGPT (800 million weekly users as of late 2025 — that's most industries now), or you're already investing in ChatGPT SEO and want every surface covered.
Skip the agonizing. Spend the hour, ship the file, then put your real budget into the table above. If you want to know whether AI engines can see your business at all — file or no file — that's exactly what our free AI visibility audit measures.
FAQ
What is llms.txt in simple terms?
It's a plain markdown file at yoursite.com/llms.txt that lists your most important pages with short descriptions, so AI systems can find and use your best content directly. Proposed by Jeremy Howard in 2024, it works like a curated table of contents for language models.
Is llms.txt the same as robots.txt?
No — they're opposites. robots.txt tells crawlers what they may not access; llms.txt tells AI systems what they should read first. robots.txt is an enforcement standard crawlers respect; llms.txt is a voluntary suggestion with no formal adoption by Google or OpenAI yet.
Does Google use llms.txt?
Not officially. Google's John Mueller has compared it to the old keywords meta tag, and no major engine has formally committed to it. Today's major AI crawlers largely ignore the file, so treat it as a low-cost bet.
How do I create an llms.txt file?
Write a markdown file: H1 with your brand name, a blockquote summary, then H2 sections listing 10–40 key pages, each with a one-sentence description. Upload it to your site root. Optionally add llms-full.txt with full page text. Then verify crawler hits in server logs monthly.
Will llms.txt get me cited in ChatGPT?
Not by itself. Does llms.txt work as a citation shortcut? No — citations come from entity strength, answer-first content, third-party mentions, and freshness; the file just makes your content easier to ingest if an agent visits. It's one small piece of generative engine optimization.
16 years in digital marketing, focused on AI SEO, GEO, AEO and local search. Every claim in this article links to its source — and every method here is what we run for real client campaigns.