What is robots.txt — blocking AI crawlers in one line, and why it can't be enforced
5 sources cross-checked
robots.txt is a text file placed at the top level of a website (for example, example.com/robots.txt) that tells search-engine and AI crawlers which paths they may fetch and which they may not. It began as a 1994 convention and became an internet standard (RFC 9309) in September 2022, but it has no enforcement power: it is a notice on the door, not a lock, so programs that ignore it are not stopped. OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User) and Anthropic (ClaudeBot and others) publish separate crawler names for training, search and user-requested fetches, so a site can block training collection while still allowing AI search to cite it. Actual blocking relies on rate limits, logins and firewalls
The three lines
Definition — a root-level text file listing allowed and disallowed paths; standardized as RFC 9309 in 2022
Limits — voluntary; can't stop bots that ignore it or AI agents that browse like people
AI era — block or allow training, search and user-requested crawlers separately by name
Key questions
How to block AI crawlers with robots.txt
**Add Disallow rules per crawler name (User-agent).**
| Goal | Rule |
|---|---|
| Block OpenAI training | `User-agent: GPTBot` / `Disallow: /` |
| Block Anthropic training | `User-agent: ClaudeBot` / `Disallow: /` |
| Opt out of Gemini training | `User-agent: Google-Extended` / `Disallow: /` |
| Block Common Crawl | `User-agent: CCBot` / `Disallow: /` |
| Keep AI search citations | Don't block `OAI-SearchBot`, `Claude-SearchBot`, etc. |
Is robots.txt legally enforceable
**It isn't a technical barrier; compliance is up to the crawler.**
| Item | robots.txt | Real enforcement |
|---|---|---|
| Nature | Notice | Lock |
| Bots that ignore it | Not stopped | Can be stopped |
| Examples | Disallow rules | Rate limits, logins, IP/firewall blocks, CDN bot management |
robots.txt vs noindex
**robots.txt says 'don't fetch'; noindex says 'don't show in search results.'**
| Item | robots.txt Disallow | noindex tag |
|---|---|---|
| Prevents | Crawling | Indexing |
| Location | Site-root file | Each page's HTML or headers |
| Caveat | Blocked URLs can still be indexed via links | Crawler must read the page to see it |
robots.txt is a "staff only" sign on the door. Polite visitors turn back; it does nothing about those who ignore it. After the Wikimedia Foundation disclosed in October 2026 that agents it believes were OpenAI's made unapproved edits and flooded its APIs (see "Wikimedia says OpenAI agents edited its wikis"), the web's oldest way of telling machines where they may go is back in focus. Here is what it can and can't do.
1. What robots.txt is
A plain text file at the root of a site (for example https://example.com/robots.txt). Crawlers read it before fetching pages and follow its rules.
Item
Detail
Location
/robots.txt at the site root
Origin
Convention proposed by Martijn Koster in 1994
Standard
IETF RFC 9309, September 2022
Audience
Crawlers and other automated fetchers
Nature
Voluntary; no authentication or access control
Directive
Meaning
Example
User-agent
Which crawler the rules address
User-agent: GPTBot, or * for all
Disallow
Paths not to fetch
Disallow: /private/; all: Disallow: /
Allow
Exception inside a disallowed path
Allow: /private/public-page
Sitemap
Sitemap location (extension outside RFC 9309)
Sitemap: https://example.com/sitemap.xml
RFC 9309 also sets practical rules: crawlers must parse at least 500 KiB, shouldn't rely on a cached copy older than 24 hours, treat a missing file (4xx) as full permission, and treat an unreachable file due to server error (5xx) as full disallow.
2. AI crawlers have several names: training, search, user requests
Company
Training collection
Search / citation
User-requested fetch
OpenAI
GPTBot
OAI-SearchBot
ChatGPT-User
Anthropic
ClaudeBot
Claude-SearchBot
Claude-User
Google
Google-Extended (opt-out token for Gemini training)
Googlebot (regular search)
—
Others
CCBot (Common Crawl), Applebot-Extended
PerplexityBot
—
The split matters because the effects differ. Block a training crawler and your content stays out of model training data. Block a search crawler and you disappear from AI answers' source lists. A common setup blocks training bots (GPTBot, ClaudeBot, Google-Extended) and allows search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot).
Setup
Result
Block training, allow search
Not used for training; can still be cited in AI search
Block both
Out of training and citations
Allow both
Default
Block Googlebot
Removed from Google Search itself — be careful
3. Why it can't stop anyone: sign versus lock
robots.txt can
robots.txt can't
Tell compliant crawlers the boundaries
Physically block programs that ignore it
Give different rules per crawler name
Identify programs that hide or fake their name
Publicly state the site's wishes
Pull back data already collected
RFC 9309 itself says the file is not a form of access control or security. Listing secret paths in it only advertises them. AI agents widen the gap: a crawler sweeps a site, while an agent opens a browser on a person's behalf and visits pages one by one, and companies' guidance on whether robots.txt applies to user-requested fetches has varied. The agents in the Wikimedia case skipped bot approval and tried to misuse a citation tool, behavior a notice can't stop.
Real enforcement
What it stops
Rate limits
Bulk requests in short periods ("What an API rate limit is")
Login and authentication
Anonymous access
IP and firewall blocks
Specific addresses or ranges
CDN bot management
Bots identified by behavior
Terms of service
Legal basis for action
4. FAQ
Question
Answer
Does Disallow remove pages from search?
No, it blocks fetching only; URLs can still appear via links. Use noindex to keep pages out of results
Does blocking undo past training?
No, it applies to future collection
Can it sit in a subfolder?
No, it must be at the root; each subdomain needs its own
Case sensitive?
Paths yes (/Private ≠ /private); directive names no
What is llms.txt?
A proposed file guiding AI to a site's key content; not a standard and not an access rule
Want to be cited by AI?
Don't block AI search crawlers, and write clearly structured, well-sourced pages
5. What remains unclear
Changing crawler names: verify the list against each company's current docs.
Agents: no industry agreement exists on how robots.txt applies to agents acting for users; Wikimedia instead asks that AI traffic be labeled.
CDN defaults: a reported Cloudflare default change in September 2026 rests on one secondary source.