Search engines discovered the web through a convention: robots.txt plus links. AI agents are slowly developing their own conventions, and llms.txt is the most practical one so far. It is a plain-text, markdown-structured summary of a site's content, written for language models to read directly.

How the pieces fit

  • robots.txt — machine instructions about crawling boundaries. It is for crawlers, and it is about what not to fetch.
  • sitemap.xml — a machine-readable list of URLs. It is about discoverability, but it tells a reader almost nothing about the content.
  • llms.txt — human-readable, model-readable context. It answers the question a crawler never asks: "What is this site, and what can I do here?"

What belongs in an llms.txt

The best examples are short and factual: a one-paragraph overview, the key workflows, a categorized list of the actual content, and the endpoints an agent would call. No marketing copy, no keyword stuffing — the reader is a model that will act on the text.

Going deeper on this theme: How to Remain Valuable When Intelligence Becomes Cheap — the scarce human, economic, and strategic advantages that stay valuable when AI does the cognitive work. A 224-page practical book, $3.84. Read it on Gumroad →

This site's llms.txt lists every tool category, explains how to fetch the manifest and execute tools, and points to the API endpoints. An agent can read it and immediately know: what the site is, how to call it, and that no authentication is required.

Why this matters now

Agents are becoming primary visitors, not an afterthought. Publishing for them is cheap — a single text file — and it changes the quality of what they do with the site. A well-written llms.txt is the difference between an agent guessing from HTML and an agent acting from a precise contract.