BeamciteSaaS

A robots.txt that treats AI crawlers on purpose.

Most robots.txt files predate AI crawlers, so GPTBot, ClaudeBot and PerplexityBot get whatever the old rules happen to say. Decide bot by bot, or use a preset, and get a clean file to copy or download. Free, instant, no sign-up.

GPTBotOpenAI

Training data for GPT models

OAI-SearchBotOpenAI

ChatGPT search index (citations) (used for live answers and citations)

ChatGPT-UserOpenAI

Live page visits when a user asks ChatGPT (used for live answers and citations)

ClaudeBotAnthropic

Crawling for Claude

Claude-UserAnthropic

Live page visits when a user asks Claude (used for live answers and citations)

PerplexityBotPerplexity

Perplexity search index (citations) (used for live answers and citations)

Perplexity-UserPerplexity

Live page visits when a user asks Perplexity (used for live answers and citations)

Google-ExtendedGoogle

Gemini training and grounding

Applebot-ExtendedApple

Apple Intelligence training

Meta-ExternalAgentMeta

Meta AI training and index

BytespiderByteDance

Doubao / ByteDance model training

CCBotCommon Crawl

Open crawl used by many AI training sets

Press Generate robots.txt to build your file

Upload the file to your site root so it is reachable at yourdomain.com/robots.txt. Blocked crawlers get their own explicit group; allowed crawlers need no group at all and simply follow the same * rules as every other well-behaved bot.

Trade-off. Blocking AI crawlers is a legitimate choice, but it has a cost: an assistant that cannot read your pages can never cite or recommend them. If AI visibility matters to your business, the usual move is to allow at least the citation and live-answer bots.

01 / How it works

Step 1

Pick a preset

Allow everything, block only the training crawlers, or block all 12 AI bots. Presets set the toggles; you can then override any single bot.

Step 2

Tune the details

Flip individual bots, add Disallow paths that apply to every crawler, and point at your sitemap.

Step 3

Copy or download

Press Generate robots.txt to build the file, then copy it or download, drop it at your site root, and verify with the AI Crawler Access Checker.

02 / FAQ

Questions, answered.

What does robots.txt actually do?+

It is a plain-text file at your site root that tells well-behaved crawlers which parts of your site they may fetch. Each 'User-agent' group names a bot and lists Allow and Disallow rules. It is a convention, not a lock: compliant crawlers honor it, malicious scrapers ignore it.

Which AI crawlers should I allow?+

That depends on what you want. Crawlers like OAI-SearchBot, PerplexityBot and the -User agents fetch pages to answer and cite live questions, so blocking them removes you from AI answers. Crawlers like GPTBot, Google-Extended and CCBot feed model training. Many sites allow the citation bots and decide separately about training. The 'Block training, keep citations' preset does exactly that.

If I block GPTBot, does ChatGPT forget my site?+

Blocking stops future crawling; it does not retroactively remove anything already in an index or a trained model. Over time your content stops being refreshed and your pages cannot be fetched live, so your presence in that assistant's answers decays.

Why do allowed AI bots not appear in the generated file?+

Because of how the robots exclusion protocol works: a crawler that finds its own User-agent group ignores the * group entirely. Giving an allowed bot its own 'Allow: /' group would accidentally exempt it from your normal Disallow rules. Leaving it out means it follows the same * rules as every other crawler, which is almost always what you want.

Does robots.txt affect my Google Search rankings?+

This generator only manages AI crawler access and your general Disallow rules; it does not touch Googlebot. Beamcite is strictly about visibility in AI assistants, not Google rankings. Be careful with blanket 'Disallow: /' rules under *, though: those block everything, including search engines.

Where do I put the finished file?+

Upload it to your web root so it is served at yourdomain.com/robots.txt, replacing any existing file. Most platforms (Next.js, WordPress, Shopify, Netlify) have a documented way to set a custom robots.txt. Then verify with our AI Crawler Access Checker.

Is my input stored anywhere?+

No. The file is assembled in your browser when you press Generate; nothing is uploaded, logged or saved.

Free strategy call · 30 minutes

Letting AI crawlers in is step one. Cited is the goal.

Beamcite makes sure the crawlers can read you, then does the ongoing work of getting you actually cited by ChatGPT, Perplexity and Gemini, with a human running it.

Book the call

No payment. No obligation. A real human.