AI crawlers come in three kinds. Search bots collect pages so AI search can show and cite them. Training bots collect data for model training. User-triggered fetchers open a page on the spot because someone asked a question. If you want to appear in AI search, keep the search bots open and decide on the training bots according to your own policy. The two are independent: you can block GPTBot and still show up in ChatGPT search, as long as OAI-SearchBot is allowed.
This guide combines the crawler documentation published by OpenAI, Anthropic, Perplexity, Google and Apple with robots.txt files we checked on Korean content platforms on October 4, 2026.
Key takeaways
- Search bots: OAI-SearchBot (ChatGPT), PerplexityBot, Claude-SearchBot, Googlebot (which also feeds AI Overviews and AI Mode), Bingbot (Copilot) and Yeti (Naver). Block one and you drop out of that engine's answers.
- Training bots and control tokens: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and others. Their operators say blocking them doesn't affect search. The exception to watch: blocking Google-Extended also keeps your pages out of grounding in the Gemini app.
- User-triggered fetchers such as ChatGPT-User, Claude-User and Perplexity-User act on a person's request, so robots.txt may not apply.
- Naver Blog, Naver Cafe and Knowledge iN block the search bots of ChatGPT, Perplexity and Claude. Brunch blocks training bots but allows AI search bots.
- CDN and firewall bot rules often block AI bots before robots.txt matters. Test the real response.
The main AI crawlers
| Operator | Name | Type | If you block it |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Search | Not shown in ChatGPT search answers, though navigational links may still appear |
| OpenAI | GPTBot | Training | Excluded from training data; separate from search |
| OpenAI | ChatGPT-User | User-triggered | robots.txt may not apply, since a user started the request |
| Anthropic | Claude-SearchBot | Search | Not indexed for Claude's search, which may reduce visibility in results |
| Anthropic | ClaudeBot | Training | Future content excluded from training |
| Anthropic | Claude-User | User-triggered | Claude can't fetch your pages when users ask |
| Perplexity | PerplexityBot | Search | Dropped from Perplexity results; Perplexity says it isn't used for training |
| Perplexity | Perplexity-User | User-triggered | Generally ignores robots.txt |
| Googlebot | Search | Out of Google Search, including AI Overviews and AI Mode | |
| Google-Extended | Control token | Excluded from Gemini training and Gemini app grounding; no effect on Google Search | |
| Microsoft | Bingbot | Search | Out of Bing and Copilot answers |
| Apple | Applebot-Extended | Control token | Excluded from Apple's generative AI training; it doesn't crawl on its own |
| Common Crawl | CCBot | Collection | Out of a public dataset widely used for AI training |
| Naver | Yeti | Search | Out of Naver's web search, so AI Briefing can't cite your site as a web page |
Google-Extended and Applebot-Extended don't crawl anything themselves. They're labels that decide how data already fetched by Googlebot and Applebot may be used, which is why we call them control tokens.
ChatGPT search uses pages collected by OAI-SearchBot and results from third-party search providers, with Bing widely reported as one of them. If ChatGPT visibility matters to you, keep Bingbot open too.
robots.txt examples by goal
Show up in AI search, stay out of training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Think separately about Google-Extended. Blocking it keeps you out of Gemini training, but also out of grounding, when the Gemini app looks things up in Google Search to answer. If visibility in the Gemini app matters, leave it open.
Open to every AI
User-agent: *
Allow: /
Disallow: /api/
Sitemap: https://example.com/sitemap.xml
This is what anymorph.ai uses. We think accurate information about a brand ending up in training data helps AI answers over time. A publisher whose content is the product may reasonably decide otherwise.
Block only some paths
User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /search
Sitemap: https://example.com/sitemap.xml
Block paths that never need to be a source, like admin screens, carts and internal search results. robots.txt is a request, not access control. Pages that must stay private need a login or noindex.
How do Korean platforms treat AI crawlers?
If you publish on outside platforms, their robots.txt decides your AI visibility there. On October 4, 2026 we downloaded each platform's robots.txt and evaluated it against a typical post URL for each crawler.
| Platform | Google Search | Bing | ChatGPT search | Perplexity | Claude search | Training bots (GPTBot, ClaudeBot, Google-Extended) |
|---|---|---|---|---|---|---|
| Naver Blog | Allowed | Allowed | Blocked | Blocked | Blocked | Blocked |
| Naver Cafe | Blocked | Blocked | Blocked | Blocked | Blocked | Blocked |
| Knowledge iN | Allowed | Allowed | Blocked | Blocked | Blocked | Blocked |
| Naver Premium Contents | Allowed | Allowed | Blocked | Blocked | Blocked | Blocked |
| Brunch | Allowed | Allowed | Allowed | Allowed | Allowed | Blocked |
| Tistory | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed |
| velog | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed |
| Daum Cafe (public boards) | Allowed | Allowed | Allowed | Allowed | Allowed | Allowed |
Naver's robots.txt files state that bot access for AI training and retrieval-augmented generation (RAG) is prohibited. So a Naver Blog post can reach Naver AI Briefing and Google Search, but rarely becomes a source for ChatGPT, Perplexity, Claude or the Gemini app. Brunch, Kakao's writing platform, blocks training bots while allowing AI search bots, a clean split between search visibility and training opt-out. The Tistory row is based on Tistory's official notice blog, and many Daum Cafe posts are members-only, so less is readable in practice.
These are the rules each platform declares. What each AI company actually collects can differ, and the rules can change at any time. For Naver strategy specifically, see how Naver AI Briefing picks its sources.
Check your CDN and firewall first
robots.txt asks crawlers to behave. CDNs and firewalls actually block requests. If your firewall blocks AI bots, an open robots.txt doesn't help.
- Cloudflare started blocking AI crawlers by default for new domains on July 1, 2025. Since July 2026, customers on every plan can allow or block AI traffic by behavior: search, agent and training.
- If Cloudflare's managed robots.txt is on, it adds rules blocking training crawlers and its content signals (search, ai-input, ai-train). The file you wrote and the file being served may differ.
- Check other CDNs' and hosts' bot protection the same way.
The quickest check is to send a request with a crawler's user agent:
curl -I -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot)" https://example.com/
A 403 or a challenge page instead of a 200 means something is blocking. Firewalls that also check IP addresses treat spoofed requests differently from real bots, though, so the final check is your server logs: look for the real bot's requests and the status codes they got. OpenAI and Perplexity publish their bots' IP ranges, and Googlebot and Yeti can be verified with a reverse DNS lookup.
Do you need llms.txt?
It's not urgent. Google's 2026 guide says Google Search doesn't use llms.txt and that creating one neither helps nor hurts your visibility. The crawler documentation of OpenAI, Anthropic and Perplexity doesn't describe using llms.txt for search or citations either.
We keep an llms.txt at anymorph.ai, but we don't expect it to raise visibility. It takes minutes and costs nothing. The priorities are elsewhere: let search bots in, keep a sitemap, put content in text, and publish pages worth citing.
FAQ
If I block GPTBot, will I disappear from ChatGPT?
No. GPTBot is for training. Visibility in ChatGPT search is controlled by OAI-SearchBot, and OpenAI says each setting is independent. To leave ChatGPT search, block OAI-SearchBot.
How fast do robots.txt changes take effect?
OpenAI says it can take about 24 hours for search results to reflect a robots.txt change. Other operators re-read robots.txt periodically too, so give it a day or two before checking.
What happens if I block Naver's Yeti?
Your pages won't be collected for Naver's web search, so they won't appear in its web results and AI Briefing can't cite your site as a web page. If you have customers in Korea, keep Yeti open.
How do I know a bot is genuine?
OpenAI publishes IP ranges for OAI-SearchBot, GPTBot and ChatGPT-User, and Perplexity does the same for PerplexityBot and Perplexity-User. Googlebot and Yeti can be verified with a reverse DNS lookup that resolves to google.com or naver.com.
Sources
- OpenAI, Overview of OpenAI Crawlers
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity, Perplexity Crawlers
- Google, List of Google's common crawlers
- Apple, About Applebot
- Common Crawl, CCBot
- Cloudflare, Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large (July 2025)
- Cloudflare, New options to manage AI traffic (July 2026)
- Google Search Central, Optimizing your website for generative AI features on Google Search
- TechTarget, ChatGPT search: Details about OpenAI's search engine
- Anymorph, robots.txt check of Korean platforms (October 4, 2026)