Free tool · robots.txt
Write a robots.txt, and see which AI crawlers it lets in.
This robots.txt generator writes the file with an explicit choice for each AI crawler, from GPTBot and ClaudeBot to PerplexityBot and Google-Extended, and the checker reads any site's robots.txt and shows, crawler by crawler, whether it may fetch a page and which line decides it. It reads the file the way RFC 9309 says crawlers do, with the same code Heddle's gates use.
User-agent: * Allow: /
What this file says to each crawler
| Crawler | For | / | Decided by |
|---|---|---|---|
| GPTBotOpenAI | Training | Allowed | Allow: / (line 2, from *) |
| OAI-SearchBotOpenAI | Search index | Allowed | Allow: / (line 2, from *) |
| ChatGPT-UserOpenAI | Fetches for a user | Allowed | Allow: / (line 2, from *) |
| ClaudeBotAnthropic | Training | Allowed | Allow: / (line 2, from *) |
| Claude-SearchBotAnthropic | Search index | Allowed | Allow: / (line 2, from *) |
| Claude-UserAnthropic | Fetches for a user | Allowed | Allow: / (line 2, from *) |
| PerplexityBotPerplexity | Search index | Allowed | Allow: / (line 2, from *) |
| Perplexity-UserPerplexity | Fetches for a user | Allowed | Allow: / (line 2, from *) |
| Google-ExtendedGoogle | Training | Allowed | Allow: / (line 2, from *) |
| Applebot-ExtendedApple | Training | Allowed | Allow: / (line 2, from *) |
| CCBotCommon Crawl | Training | Allowed | Allow: / (line 2, from *) |
| meta-externalagentMeta | Training | Allowed | Allow: / (line 2, from *) |
| BytespiderByteDance | Training | Allowed | Allow: / (line 2, from *) |
The file itself
- Warningno Sitemap line; crawlers find the sitemap faster when robots.txt names it
How the tool decides
Rules are matched the way RFC 9309 says crawlers match them. A crawler obeys the groups that name it and ignores the * group if any do; within its rules the longest matching path wins, Allow wins a tie, * matches any run of characters and a trailing $ anchors the end. Lines before any User-agent, unknown directives, relative Sitemap lines and files over Google's 500 KiB limit are reported.
Training, search and user crawlers
The vendors now split their crawlers by job, and the split is the decision you're making. GPTBot and ClaudeBot collect pages that may be used to train models. OAI-SearchBot, Claude-SearchBot and PerplexityBot build the index those assistants cite from, so blocking them keeps you out of their answers. ChatGPT-User, Claude-User and Perplexity-User fetch a page because a person asked; OpenAI and Perplexity say these may not follow robots.txt at all.
Google-Extended isn't a crawler. It's a token that says whether pages Google already fetched may be used for Gemini, and Google says it doesn't affect inclusion or ranking in Search. Blocking training while allowing search is a common choice, and the generator makes it one click.
What the checker can't tell you
A robots.txt is a request. The checker reports what your file asks of each crawler, not what each crawler does, and it doesn't see a firewall or CDN rule that blocks a bot before it reaches the file. If your host or CDN has a setting for AI bots, check that too.
On the sites Heddle runs, the agent-ready gate fails a build that blocks the answer-engine crawlers, unless the company has written down that it means to.
Sources
- rfc-editor.org/rfc/rfc9309.html
- developers.google.com/crawling/docs/robots-txt/robots-txt-spec
- developers.openai.com/api/docs/bots
- support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- docs.perplexity.ai/docs/resources/perplexity-crawlers
- developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
Each was read on October 11, 2026.
Questions
Does blocking GPTBot keep me out of ChatGPT search?
No. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the one that surfaces sites in ChatGPT search, so they're separate choices.
Will this change how Google ranks me?
Only if you block Googlebot, which the generator never does. Google-Extended controls use for Gemini, and Google says it isn't a ranking signal.
Does anything I type get stored?
No. The generator and checker run in your browser. To check a live file, our server fetches only /robots.txt from the address you give and passes it back; it keeps nothing.
Work with us
Have the agents keep it true.
A file you write once goes stale the next time the site changes. Heddle writes llms.txt, robots.txt and the markup on every deploy, and its gates fail the build when one of them is wrong. Tell us about your site and we'll show you what they would catch.
On the call, for your site
- The searches your buyers makeMeasured, with how many people make each one
- Where you show up, and where you don'tOn Google and in ChatGPT's answers
- The agents' first monthThe pages, fixes and links they would start with