Blog · Guide

AI Crawlers and robots.txt: Which Bots to Allow for AI Search

Blocking GPTBot doesn't keep you out of ChatGPT search. How to tell AI search bots from training bots, and what Korean platforms like Naver Blog allow.

Anymorph Team7 min read

AI crawlers come in three kinds. Search bots collect pages so AI search can show and cite them. Training bots collect data for model training. User-triggered fetchers open a page on the spot because someone asked a question. If you want to appear in AI search, keep the search bots open and decide on the training bots according to your own policy. The two are independent: you can block GPTBot and still show up in ChatGPT search, as long as OAI-SearchBot is allowed.

This guide combines the crawler documentation published by OpenAI, Anthropic, Perplexity, Google and Apple with robots.txt files we checked on Korean content platforms on October 4, 2026.

Key takeaways

  • Search bots: OAI-SearchBot (ChatGPT), PerplexityBot, Claude-SearchBot, Googlebot (which also feeds AI Overviews and AI Mode), Bingbot (Copilot) and Yeti (Naver). Block one and you drop out of that engine's answers.
  • Training bots and control tokens: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and others. Their operators say blocking them doesn't affect search. The exception to watch: blocking Google-Extended also keeps your pages out of grounding in the Gemini app.
  • User-triggered fetchers such as ChatGPT-User, Claude-User and Perplexity-User act on a person's request, so robots.txt may not apply.
  • Naver Blog, Naver Cafe and Knowledge iN block the search bots of ChatGPT, Perplexity and Claude. Brunch blocks training bots but allows AI search bots.
  • CDN and firewall bot rules often block AI bots before robots.txt matters. Test the real response.

The main AI crawlers

OperatorNameTypeIf you block it
OpenAIOAI-SearchBotSearchNot shown in ChatGPT search answers, though navigational links may still appear
OpenAIGPTBotTrainingExcluded from training data; separate from search
OpenAIChatGPT-UserUser-triggeredrobots.txt may not apply, since a user started the request
AnthropicClaude-SearchBotSearchNot indexed for Claude's search, which may reduce visibility in results
AnthropicClaudeBotTrainingFuture content excluded from training
AnthropicClaude-UserUser-triggeredClaude can't fetch your pages when users ask
PerplexityPerplexityBotSearchDropped from Perplexity results; Perplexity says it isn't used for training
PerplexityPerplexity-UserUser-triggeredGenerally ignores robots.txt
GoogleGooglebotSearchOut of Google Search, including AI Overviews and AI Mode
GoogleGoogle-ExtendedControl tokenExcluded from Gemini training and Gemini app grounding; no effect on Google Search
MicrosoftBingbotSearchOut of Bing and Copilot answers
AppleApplebot-ExtendedControl tokenExcluded from Apple's generative AI training; it doesn't crawl on its own
Common CrawlCCBotCollectionOut of a public dataset widely used for AI training
NaverYetiSearchOut of Naver's web search, so AI Briefing can't cite your site as a web page

Google-Extended and Applebot-Extended don't crawl anything themselves. They're labels that decide how data already fetched by Googlebot and Applebot may be used, which is why we call them control tokens.

ChatGPT search uses pages collected by OAI-SearchBot and results from third-party search providers, with Bing widely reported as one of them. If ChatGPT visibility matters to you, keep Bingbot open too.

robots.txt examples by goal

Show up in AI search, stay out of training

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

Think separately about Google-Extended. Blocking it keeps you out of Gemini training, but also out of grounding, when the Gemini app looks things up in Google Search to answer. If visibility in the Gemini app matters, leave it open.

Open to every AI

User-agent: *
Allow: /
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

This is what anymorph.ai uses. We think accurate information about a brand ending up in training data helps AI answers over time. A publisher whose content is the product may reasonably decide otherwise.

Block only some paths

User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /search

Sitemap: https://example.com/sitemap.xml

Block paths that never need to be a source, like admin screens, carts and internal search results. robots.txt is a request, not access control. Pages that must stay private need a login or noindex.

How do Korean platforms treat AI crawlers?

If you publish on outside platforms, their robots.txt decides your AI visibility there. On October 4, 2026 we downloaded each platform's robots.txt and evaluated it against a typical post URL for each crawler.

PlatformGoogle SearchBingChatGPT searchPerplexityClaude searchTraining bots (GPTBot, ClaudeBot, Google-Extended)
Naver BlogAllowedAllowedBlockedBlockedBlockedBlocked
Naver CafeBlockedBlockedBlockedBlockedBlockedBlocked
Knowledge iNAllowedAllowedBlockedBlockedBlockedBlocked
Naver Premium ContentsAllowedAllowedBlockedBlockedBlockedBlocked
BrunchAllowedAllowedAllowedAllowedAllowedBlocked
TistoryAllowedAllowedAllowedAllowedAllowedAllowed
velogAllowedAllowedAllowedAllowedAllowedAllowed
Daum Cafe (public boards)AllowedAllowedAllowedAllowedAllowedAllowed

Naver's robots.txt files state that bot access for AI training and retrieval-augmented generation (RAG) is prohibited. So a Naver Blog post can reach Naver AI Briefing and Google Search, but rarely becomes a source for ChatGPT, Perplexity, Claude or the Gemini app. Brunch, Kakao's writing platform, blocks training bots while allowing AI search bots, a clean split between search visibility and training opt-out. The Tistory row is based on Tistory's official notice blog, and many Daum Cafe posts are members-only, so less is readable in practice.

These are the rules each platform declares. What each AI company actually collects can differ, and the rules can change at any time. For Naver strategy specifically, see how Naver AI Briefing picks its sources.

Check your CDN and firewall first

robots.txt asks crawlers to behave. CDNs and firewalls actually block requests. If your firewall blocks AI bots, an open robots.txt doesn't help.

  • Cloudflare started blocking AI crawlers by default for new domains on July 1, 2025. Since July 2026, customers on every plan can allow or block AI traffic by behavior: search, agent and training.
  • If Cloudflare's managed robots.txt is on, it adds rules blocking training crawlers and its content signals (search, ai-input, ai-train). The file you wrote and the file being served may differ.
  • Check other CDNs' and hosts' bot protection the same way.

The quickest check is to send a request with a crawler's user agent:

curl -I -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot)" https://example.com/

A 403 or a challenge page instead of a 200 means something is blocking. Firewalls that also check IP addresses treat spoofed requests differently from real bots, though, so the final check is your server logs: look for the real bot's requests and the status codes they got. OpenAI and Perplexity publish their bots' IP ranges, and Googlebot and Yeti can be verified with a reverse DNS lookup.

Do you need llms.txt?

It's not urgent. Google's 2026 guide says Google Search doesn't use llms.txt and that creating one neither helps nor hurts your visibility. The crawler documentation of OpenAI, Anthropic and Perplexity doesn't describe using llms.txt for search or citations either.

We keep an llms.txt at anymorph.ai, but we don't expect it to raise visibility. It takes minutes and costs nothing. The priorities are elsewhere: let search bots in, keep a sitemap, put content in text, and publish pages worth citing.

FAQ

If I block GPTBot, will I disappear from ChatGPT?

No. GPTBot is for training. Visibility in ChatGPT search is controlled by OAI-SearchBot, and OpenAI says each setting is independent. To leave ChatGPT search, block OAI-SearchBot.

How fast do robots.txt changes take effect?

OpenAI says it can take about 24 hours for search results to reflect a robots.txt change. Other operators re-read robots.txt periodically too, so give it a day or two before checking.

What happens if I block Naver's Yeti?

Your pages won't be collected for Naver's web search, so they won't appear in its web results and AI Briefing can't cite your site as a web page. If you have customers in Korea, keep Yeti open.

How do I know a bot is genuine?

OpenAI publishes IP ranges for OAI-SearchBot, GPTBot and ChatGPT-User, and Perplexity does the same for PerplexityBot and Perplexity-User. Googlebot and Yeti can be verified with a reverse DNS lookup that resolves to google.com or naver.com.

Sources

What is AI saying about your brand right now?

Get a product demo and a report on where your AI visibility stands today.