At the root of nearly every website sits a plain-text file called robots.txt — often fewer than twenty lines — that instructs search engine crawlers where they're welcome. It's simple, powerful, and dangerously easy to get wrong.

How it works

Before crawling a site, well-behaved bots fetch /robots.txt and read its rules. Each rule block names a crawler (User-agent) and lists paths to avoid (Disallow) or explicitly permit (Allow). A typical file:

Related reading: How to Remain Valuable When Intelligence Becomes Cheap — a 224-page practical book on staying valuable as intelligence gets cheap. $3.84. Read it on Gumroad →

User-agent: *
Disallow: /admin/
Disallow: /tmp/
Allow: /

Sitemap: https://example.com/sitemap.xml

This says: all crawlers, stay out of /admin/ and /tmp/, everything else is fine, and here's the sitemap. For background on what crawlers do with these instructions, see how search engines crawl and index websites.

What robots.txt cannot do

Three misconceptions cause most robots.txt incidents. First, it's advisory, not access control — malicious bots ignore it, and anything disallowed is still publicly reachable by URL. Never use it to "hide" sensitive pages; use authentication. Second, disallowed pages can still appear in search results — Google may index a blocked URL's existence (without its content) if other pages link to it. Third, it's public by design — attackers read robots.txt files to discover your /admin/ and /backup/ paths, so don't advertise locations you haven't secured.

Common mistakes

  • Blocking everything in production. A staging Disallow: / copied to the live site de-indexes the whole domain. This happens constantly — check after every deploy.
  • Blocking CSS and JS. Google needs your stylesheets and scripts to render pages. Blocking /assets/ can tank rankings because the crawler sees a broken page.
  • Trailing-slash confusion. Disallow: /blog blocks /blog, /blog/, and /blogroll. Paths are prefix matches — be precise.

Test before you ship

Rules are cheap to verify and expensive to get wrong. The robots.txt tester checks whether specific URLs are crawlable under your rules, and the robots.txt analyser reviews the file itself for issues. Test every change against your most important URLs — homepage, key landing pages, sitemap — before it goes live.