---
title: 'How to write an llms.txt that AI crawlers will actually use (with examples)'
url: 'https://maw11.preview.mountainairweb.com/blog/how-to-write-an-llms-txt'
markdown: 'https://maw11.preview.mountainairweb.com/blog/how-to-write-an-llms-txt.md'
date: '2026-09-23'
description: 'Short answer: An llms.txt is a plain Markdown file at your site root: one H1 with your name, one blockquote that says what you are and where, then H2 sections of links with a one-line note each, pointing at Markdown versions of your pages wherever you can serve them. Keep it short, hand-written, an…'
taxonomy:
  category:
    - 'AI & Search'
  tag:
    - llms.txt
    - 'AI crawlers'
    - robots.txt
---

[AI & Search](https://maw11.preview.mountainairweb.com/blog/category:AI%20%26%20Search) September 23, 2026 12 min read 

# How to write an llms.txt that AI crawlers will actually use (with examples)

The file takes twenty minutes. The part that matters is everything you have to do around it.

  Nicholas Murray Author  

**Short answer:** An `llms.txt` is a plain Markdown file at your site root: one H1 with your name, one blockquote that says what you are and where, then H2 sections of links with a one-line note each, pointing at Markdown versions of your pages wherever you can serve them. Keep it short, hand-written, and regenerated when pages change. Be clear-eyed about the return: in May 2026 Ahrefs found 97% of `llms.txt` files on 137,000 domains received no requests at all, and a 12-week log study across 83 sites counted 7 fetches from OpenAI and 9 from Anthropic against thousands of `robots.txt` reads. No AI search engine has said it uses the file. It still costs nothing, coding agents do read it, and writing it forces you to produce the one-paragraph entity statement every other AI signal depends on. The useful version of the job is the whole layer at once: `llms.txt`, per-crawler `robots.txt` rules, Markdown routes, entity schema, and logs that show whether any of it is being read. That bundle is what our [MawGeo](https://maw11.preview.mountainairweb.com/services/mawgeo) product installs; the file on its own is the cheapest and least consequential piece.

## What the llms.txt standard actually asks for

The proposal comes from Jeremy Howard of Answer.AI, published in September 2024 and revised in August 2026 ([llmstxt.org](https://llmstxt.org/)). It is deliberately small. The spec requires an H1 with the site or project name. Everything else is optional: a blockquote with a short summary, a few paragraphs of context, then H2 headings each followed by a Markdown list of `[name](url): note` links. An H2 called `Optional` marks links a reader can skip when its context window is tight. The spec also proposes that every page be available as Markdown by appending `.md` to its URL, advertised with `rel="alternate"` links, so an agent can fetch clean text instead of a 300 KB HTML page.

`llms-full.txt` is not in the spec. It is a documentation-platform convention: the whole site concatenated into one Markdown file for a single request. Anthropic's developer docs do both, with an index of roughly 700 `.md` links and a pointer to `llms-full.txt` at the end ([platform.claude.com](https://platform.claude.com/docs/llms.txt)). Stripe added its file in March 2025 with an instructions section that tells assistants to prefer the Checkout Sessions API and never recommend the deprecated Card Element ([Apideck](https://www.apideck.com/blog/stripe-llms-txt-instructions-section)). Cloudflare, Vercel and Supabase publish them too. Every well-known example is a developer-documentation site whose readers are coding agents. That is the audience with evidence behind it.

## Does anyone fetch it?

Ahrefs analyzed 137,210 domains in May 2026: 28% publish an `llms.txt`, 97% of those files received zero requests that month, and 77% of the requests that did arrive came from non-AI bots, mostly SEO audit tools. GPTBot was the busiest AI fetcher at 4.5% of requests, and Claude-Code, the coding agent, fetched it more often than any AI search or assistant bot ([Ahrefs](https://ahrefs.com/blog/llmstxt-study/)).

Something Inc. watched 83 sites for twelve weeks to 19 July 2026 and compared `llms.txt` fetches to `robots.txt` fetches by the same crawlers. OpenAI: 7 against 3,990. Anthropic: 9 against 3,120. PerplexityBot: 0 against 775. Meta's crawler was the outlier at 193 against 172 ([Something Inc.](https://somethinginc.com/blog/llms-txt-ai-crawlers-fetch-data/)). Three of the four major crawlers treat the file as effectively invisible.

Google's John Mueller wrote "FWIW no AI system currently uses llms.txt" in June 2025 ([Search Engine Roundtable](https://www.seroundtable.com/google-ai-llms-txt-39607.html)), and in June 2026 called its benefit "purely speculative for now," adding the line that belongs on every agency's wall: "the most basic agentic optimization is in place, namely: don't block agents. I think that hurdle will be the biggest, for most sites" ([Search Engine Journal](https://www.searchenginejournal.com/google-says-llms-txt-is-purely-speculative-for-now/577576/)).

Common Crawl's July 2026 archive adds the quality picture. Of 584,107 files, 49.9% had the full spec shape, 22.6% contained no links at all, 40.4% of "successful" responses were HTML from a catch-all route, and 10.6% were `robots.txt` served under the wrong name. Two-thirds came from plugins, Wix alone 41%, and the median All in One SEO file ran to 8,772 tokens ([Common Crawl](https://commoncrawl.org/blog/a-content-analysis-of-llms-txt-files-from-the-july-2026-crawl-archive)). Most of the web's `llms.txt` files are generated junk nobody reads.

So why write one? It is twenty minutes. Coding agents do read it, which matters if your buyers' staff research vendors in Claude Code or Cursor. And the exercise produces the entity statement, the one paragraph saying who you are, where and for whom, that you then reuse in your Organization schema, your Google Business Profile and your homepage. The paragraph is the asset. The file is where it lives.

## What each AI crawler reads and respects

Each vendor runs separate bots for training, for search indexing, and for fetching a page a user asked about, and they behave differently.

| Crawler (vendor) | Purpose | robots.txt | llms.txt | JSON-LD | Rendered HTML | Markdown |
|---|---|---|---|---|---|---|
| GPTBot (OpenAI) | Model training | Respects | Rare (7 fetches in 12 weeks) | No evidence | Yes | No evidence |
| OAI-SearchBot (OpenAI) | ChatGPT search index | Respects | No evidence | No evidence | Yes | No evidence |
| ChatGPT-User (OpenAI) | User-requested fetch | "May not apply" | No evidence | Ignored in live fetch | Visible text only | No evidence |
| ClaudeBot (Anthropic) | Model training | Respects | Rare (9 in 12 weeks) | No evidence | Yes | No evidence |
| Claude-SearchBot (Anthropic) | Search quality | Respects | No evidence | No evidence | Yes | No evidence |
| Claude-User (Anthropic) | User-requested fetch | Respects | No evidence | Ignored in live fetch | Visible text only | No evidence |
| Claude Code and similar agents | Developer tooling | Varies | Yes, top AI fetcher | No | Yes | Requests it (`Accept: text/markdown`) |
| PerplexityBot | Perplexity index, no training | Respects | 0 in 12 weeks | No evidence | Yes | No evidence |
| Perplexity-User | User-requested fetch | Generally ignores | No evidence | Ignored in live fetch | Visible text only | No evidence |
| Googlebot / Google-Extended | Search and AI Overviews / Gemini training and grounding | Respects | No | Yes, including JS-injected | Yes | No |
| Meta-ExternalAgent | Training | Respects | Yes (193 in 12 weeks) | No evidence | Yes | No evidence |

Sources: [OpenAI](https://developers.openai.com/api/docs/bots), [Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-the-web-what-is-claude-bot), [Perplexity](https://docs.perplexity.ai/guides/bots), [Google crawlers](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers), [Google structured data](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data), fetch counts from [Something Inc.](https://somethinginc.com/blog/llms-txt-ai-crawlers-fetch-data/) and [Ahrefs](https://ahrefs.com/blog/llmstxt-study/), the JSON-LD live-fetch finding from the searchVIU experiment reported by [Ahrefs](https://ahrefs.com/blog/schema-ai-citations/), Markdown behaviour from [Vercel](https://vercel.com/kb/guide/make-your-documentation-readable-by-ai-agents). "No evidence" means no vendor statement and no log study, not proof of absence.

Two rows deserve a second look. Google-Extended is a control token, not a crawler: it appears in no server log, and disallowing it stops Gemini training and grounding while Google states it "does not impact a site's inclusion in Google Search." Apple's Applebot-Extended works the same way ([Apple](https://support.apple.com/en-us/119829)). And the JSON-LD column: Google reads it and builds its knowledge graph from it; the chat engines, during a live fetch, extract the visible HTML and skip it. Schema is for being the right entity when Google describes you, not a lever for being quoted by ChatGPT.

## A full example for a multi-location business

Here is the file we would write for a fictional physiotherapy group with three clinics. Numbered notes follow.

```markdown
# Northgate Physiotherapy

> Physiotherapy group with three clinics on Vancouver Island, BC, Canada
> (Campbell River, Courtenay, Nanaimo), founded 2011, ICBC and WorkSafeBC
> claims accepted, direct billing to most extended health plans.
> Every page on this site is also available as Markdown by appending `.md`.

Clinics are owned and operated by Northgate Physiotherapy Ltd. Hours, phone
numbers and practitioner lists on each clinic page are authoritative and are
updated whenever they change.

## Clinics

- [Campbell River](https://northgatephysio.example/clinics/campbell-river.md): 1180 Shoppers Row, hours, 7 physiotherapists, parking, bus routes.
- [Courtenay](https://northgatephysio.example/clinics/courtenay.md): 2270 Cliffe Ave, hours, 5 physiotherapists, pelvic health and vestibular programs.
- [Nanaimo](https://northgatephysio.example/clinics/nanaimo.md): 6581 Aulds Rd, hours, 9 physiotherapists, open Saturdays.

## Services

- [ICBC claims after a car accident](https://northgatephysio.example/services/icbc.md): what is covered, how many sessions are pre-approved, what to bring.
- [WorkSafeBC injuries](https://northgatephysio.example/services/worksafebc.md): referral process, reporting, return-to-work plans.
- [Post-surgical rehabilitation](https://northgatephysio.example/services/post-surgical.md): knee, hip, shoulder; typical timelines.
- [Pelvic health](https://northgatephysio.example/services/pelvic-health.md): Courtenay and Nanaimo only.

## Pricing and billing

- [Fees and direct billing](https://northgatephysio.example/fees.md): fee table, which insurers bill directly, cancellation policy.

## Answers

- [How long does ICBC physiotherapy coverage last?](https://northgatephysio.example/answers/icbc-coverage-length.md)
- [Do I need a doctor's referral?](https://northgatephysio.example/answers/referral.md)
- [What to expect at a first appointment](https://northgatephysio.example/answers/first-appointment.md)

## Optional

- [About and practitioners](https://northgatephysio.example/about.md)
- [Privacy policy](https://northgatephysio.example/privacy.md)
- [Journal](https://northgatephysio.example/journal.md): clinic news, roughly monthly.
```

1. **The blockquote is the entity statement.** Category, locations, founding year, the two things a buyer asks first (ICBC, direct billing). Reuse these exact words in the Organization schema `description`, the Google Business Profile and the homepage. Consistency across sources is the signal that survives every study.
2. **One link per location, with the note carrying the facts** an answer engine needs: address, hours, headcount, what is distinctive. Each clinic page also carries its own LocalBusiness schema linked to the parent Organization, which is what Google's guidance asks for.
3. **Links go to `.md` versions**, with the HTML page one extension away. If you cannot serve Markdown, link the HTML pages; a correct HTML link beats a broken `.md` link.
4. **The Answers section is the pages you want quoted.** Question-phrased titles, answer in the first 150 words on the page itself. The file only points; the page has to deliver.
5. **Optional means skippable**, not "everything else." Thirty links, not three hundred. If an agent reads 8,000 tokens to find your phone number, you have reproduced your sitemap.
6. **No instructions to the model.** Stripe can steer coding assistants because its readers are coding assistants. A clinic telling ChatGPT it is "the best physiotherapy on Vancouver Island" is noise; Mueller's point is that self-reported files cannot help a system choose you over the clinic next door ([SEJ](https://www.searchenginejournal.com/googles-mueller-says-llms-txt-cant-help-llms-differentiate-sites/566027/)).

## What the llms.txt generators get wrong

The failure modes in the Common Crawl corpus repeat across the popular plugins and online generators.

- **Dumping the sitemap.** Most WordPress plugins list every post, product and category with its meta description: the 8,772-token median, no curation, no Optional section, nothing an agent can use faster than the sitemap it already had.
- **Serving HTML.** Single-page apps and some builders return the home page for `/llms.txt` with a 200 status; four in ten "successful" files were this. Check yours with `curl -I` for `Content-Type: text/plain` or `text/markdown`.
- **A copy of robots.txt** under the wrong name, 10.6% of the corpus.
- **Builder boilerplate** telling agents to stop scraping and call the platform's API instead, which Common Crawl found to be the dominant use of the file today.
- **Generated once, never updated.** The file says the Courtenay clinic opens at 7; the page says 8. Even the study kindest to `llms.txt` said publish it only if you keep it current.
- **AI-written summaries** of your own site. Fluent, and frequently wrong about hours, locations and what you do not offer. Write the blockquote yourself; it is four lines.

Use a generator for the link list, then edit by hand. If it cannot regenerate when a page changes, keep it short enough that a human will.

## The robots.txt rules that matter more

Crawlers read `robots.txt` thousands of times for every look at `llms.txt`, and Mueller's biggest hurdle is sites that block the agents they want to be cited by. Our rule set for a business that wants AI search visibility and is undecided on training:

```text
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

# Training bots: your decision. Blocking these does not remove you from AI search answers.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
Disallow: /

# Cloudflare Content Signals, honoured by cooperating crawlers
Content-Signal: search=yes, ai-input=yes, ai-train=no

Sitemap: https://northgatephysio.example/sitemap.xml
```

Blocking `GPTBot` leaves you in ChatGPT search; blocking `OAI-SearchBot` removes you from its answers within about 24 hours ([OpenAI](https://developers.openai.com/api/docs/bots)). PerplexityBot does not train models, so there is no trade-off there. The `Content-Signal` line is Cloudflare's 2025 extension, a preference (`search`, `ai-input`, `ai-train`) now the default on more than 3.8 million domains using its managed `robots.txt` ([Cloudflare](https://blog.cloudflare.com/content-signals-policy/)). Cloudflare's own documentation says compliance is voluntary: a sign on the door, not a lock.

## Serve Markdown, and the file starts earning its keep

The half of the spec people skip is the `.md` routes, and it is the half with usage behind it. Vercel's agent study found agents requested Markdown in 65.1% of fetches and chose it 95.7% of the time when they expressed any format preference; Claude Code sends `Accept: text/markdown` natively ([Vercel](https://vercel.com/kb/guide/make-your-documentation-readable-by-ai-agents)). The pattern: answer `Accept: text/markdown` with `Content-Type: text/markdown` and a `Vary: Accept` header, offer the same thing at `/page.md` for clients that cannot set headers, and advertise it with `<link rel="alternate" type="text/markdown">`. WorkOS documented the trap: a routing bug served the wrong page's Markdown with a 200 status and no agent could tell ([WorkOS](https://workos.com/blog/your-docs-have-a-new-audience)). Test the routes like redirects.

On WordPress this is a plugin and a rewrite rule. On the [flat-file sites](https://maw11.preview.mountainairweb.com/services/flat-file-sites) we build the content is already Markdown with YAML on top, so every page has a `.md` twin by default and the `llms.txt` links to them without a second system. Either way, the Markdown view is the fastest honest audit of a page: if the `.md` version does not state the answer in its first paragraph, neither does the HTML, whatever the design says.

## Do and don't

**Do**

- Write the H1 and blockquote by hand; reuse the blockquote as your entity description everywhere.
- Link 20 to 60 pages, grouped by what a buyer asks, with a factual note on each.
- Point at `.md` routes where they exist and at HTML where they do not.
- Serve it as `text/plain` or `text/markdown` at the root; confirm with `curl`.
- Regenerate it when a page's title, URL or key facts change.
- Log fetches of `/llms.txt` and `/*.md` by user agent.

**Don't**

- List every URL. That is the sitemap.
- Write marketing claims or instructions to the model.
- Let an SPA return HTML for it.
- Trust an AI-written summary of your own hours and services.
- Block the search bots while worrying about training bots.
- Report `llms.txt` to a client or a boss as "AI optimisation done."

## Why the file only makes sense as part of a layer

What the evidence rewards: not being blocked, pages an engine can parse, a consistent entity across your site and Google's graph, and the ability to see what the crawlers did. `llms.txt` touches none of those by itself. It is a map, useful only when the roads exist and someone checks the traffic.

That is the design behind [MawGeo](https://maw11.preview.mountainairweb.com/services/mawgeo). It ships with every site we build and installs on existing WordPress sites: Organization, LocalBusiness, Service and Article schema generated from the content and kept consistent across every location page; `robots.txt` tuned per crawler along the lines above; `llms.txt` maintained from the same page data; Markdown delivery on flat-file builds; and a dashboard that logs every AI user agent server-side and catches click-throughs from ChatGPT, Perplexity, Claude and Gemini referrers. The last piece is the one we would keep if we could keep only one, because it is the difference between a theory and a number. In the Ahrefs sample, 97% of sites that published the file could have learned from their logs that nobody read it. Most never looked.

## What to do next

Write the four-line blockquote today; it is the hardest and most reusable part. Then check `robots.txt` for the three search bots, test `/llms.txt` with `curl -I`, and ask whether a single AI user agent has appeared in your logs this month. If you would rather have that checked for you, the [free audit](https://maw11.preview.mountainairweb.com/free-audit) reports on crawler rules, entity schema, Markdown readiness and whether your logs show any AI visitor at all, with the fixes in priority order.

## More from the journal

 [See all posts](https://maw11.preview.mountainairweb.com/blog) 

AI & SearchOct 7, 2026

### What AI search optimization actually means (and what it costs to ignore)

GEO, AEO and AI visibility are the same job with different labels. Here is how ChatGPT, Perplexity, Google…

[What AI search optimization actually means (and what it costs to ignore)](https://maw11.preview.mountainairweb.com/blog/what-ai-search-optimization-means)

AI & SearchSep 30, 2026

### How to choose an AI search optimization agency: 12 questions that expose the pretenders

A buyer's guide for marketing directors and agency VPs evaluating AI search optimization services: twelve questions on measurement,…

[How to choose an AI search optimization agency: 12 questions that expose the pr…](https://maw11.preview.mountainairweb.com/blog/how-to-choose-an-ai-search-optimization-agency)

AI & SearchAug 19, 2026

### What actually gets a business cited by ChatGPT, Perplexity and Google AI

The 2026 studies are in: schema alone doesn’t move AI citations, llms.txt is barely fetched, and what works…

[What actually gets a business cited by ChatGPT, Perplexity and Google AI](https://maw11.preview.mountainairweb.com/blog/what-gets-a-business-cited-by-ai)

---

## Navigation

- Parent: [Blog](https://maw11.preview.mountainairweb.com/blog.md)
- Previous: [How to choose an AI search optimization agency: 12 questions that expose the pretenders](https://maw11.preview.mountainairweb.com/blog/how-to-choose-an-ai-search-optimization-agency.md)
- Next: [Headless WordPress in 2026: why we usually say no (and what we recommend instead)](https://maw11.preview.mountainairweb.com/blog/headless-wordpress-in-2026.md)
