
{"id":1147,"date":"2026-07-10T02:45:00","date_gmt":"2026-07-10T02:45:00","guid":{"rendered":"https:\/\/maxaeo.ai\/blog\/gptbot-claudebot-perplexitybot\/"},"modified":"2026-07-10T02:45:00","modified_gmt":"2026-07-10T02:45:00","slug":"gptbot-claudebot-perplexitybot","status":"publish","type":"post","link":"https:\/\/maxaeo.ai\/blog\/gptbot-claudebot-perplexitybot\/","title":{"rendered":"GPTBot ClaudeBot PerplexityBot Compared: AI Crawler Rules for Brand Visibility"},"content":{"rendered":"<p>If you searched for <strong>GPTBot ClaudeBot PerplexityBot<\/strong>, the real question is probably not &quot;what are these bots?&quot; It is &quot;which ones should I allow, block, verify, and monitor if I care about appearing in AI answers?&quot;<\/p>\n<p><strong>Short answer:<\/strong> <code>PerplexityBot<\/code>, <code>OAI-SearchBot<\/code>, <code>Claude-SearchBot<\/code>, and user-triggered fetchers usually matter most for current AI search visibility. <code>GPTBot<\/code> and <code>ClaudeBot<\/code> matter more for training permissions and long-term model knowledge. Treat them as different controls, not one generic &quot;AI bot&quot; setting.<\/p>\n<p>For B2B SaaS, ecommerce, publishers, and technical brands, the right policy is rarely &quot;allow every AI crawler&quot; or &quot;block every AI crawler.&quot; Segment by <strong>bot purpose<\/strong>, <strong>page type<\/strong>, <strong>verification confidence<\/strong>, and <strong>business risk<\/strong>. Then measure whether AI systems mention, cite, rank, and describe the brand accurately.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" style=\"max-width:100%;height:auto\" loading=\"lazy\"  src=\"https:\/\/maxaeo.ai\/blog\/wp-content\/uploads\/2026\/07\/1783607405584-13-5597-1.jpg\" alt=\"GPTBot ClaudeBot PerplexityBot comparison dashboard showing crawler log hits, verified IPs, and downstream AI citations\"><\/figure>\n<h2>GPTBot, ClaudeBot, and PerplexityBot in 50 Words<\/h2>\n<p><code>GPTBot<\/code>, <code>ClaudeBot<\/code>, and <code>PerplexityBot<\/code> are AI crawler user agents used by OpenAI, Anthropic, and Perplexity. They do not serve the same function: GPTBot and ClaudeBot are mainly training-related crawlers, while PerplexityBot is tied to search result surfacing and citations inside Perplexity.<\/p>\n<p>That difference changes your robots.txt policy. Blocking <code>GPTBot<\/code> is not the same as blocking ChatGPT search. Blocking <code>ClaudeBot<\/code> is not the same as blocking Claude&#39;s search and user-directed retrieval. Blocking <code>PerplexityBot<\/code> can directly reduce Perplexity&#39;s ability to fetch and cite your public pages.<\/p>\n<h2>Quick Comparison: Which Bot Should You Care About First?<\/h2>\n<p>For current AI visibility, prioritize crawlers and fetchers that feed live search and answer retrieval. For content licensing, privacy, and training opt-out decisions, prioritize training crawlers.<\/p>\n<table>\n<thead>\n<tr>\n<th>Crawler or user agent<\/th>\n<th>Owner<\/th>\n<th>Primary purpose<\/th>\n<th align=\"right\">Immediate AI visibility impact<\/th>\n<th>Policy priority<\/th>\n<th>What to monitor<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>GPTBot<\/code><\/td>\n<td>OpenAI<\/td>\n<td>Training data for OpenAI generative AI foundation models<\/td>\n<td align=\"right\">Medium to low<\/td>\n<td>Training governance<\/td>\n<td>Hits to evergreen public pages, sensitive paths, crawl rate<\/td>\n<\/tr>\n<tr>\n<td><code>OAI-SearchBot<\/code><\/td>\n<td>OpenAI<\/td>\n<td>Surfacing websites in ChatGPT search features<\/td>\n<td align=\"right\">High<\/td>\n<td>ChatGPT search visibility<\/td>\n<td>2xx access, important URL coverage, robots changes<\/td>\n<\/tr>\n<tr>\n<td><code>ChatGPT-User<\/code><\/td>\n<td>OpenAI<\/td>\n<td>User-triggered page visits from ChatGPT and Custom GPTs<\/td>\n<td align=\"right\">High for live user prompts<\/td>\n<td>Buyer-journey monitoring<\/td>\n<td>Requested URLs, WAF treatment, conversion-page access<\/td>\n<\/tr>\n<tr>\n<td><code>ClaudeBot<\/code><\/td>\n<td>Anthropic<\/td>\n<td>Training-related public web collection<\/td>\n<td align=\"right\">Medium to low<\/td>\n<td>Training governance<\/td>\n<td>Crawl frequency, crawl delay, sensitive directory access<\/td>\n<\/tr>\n<tr>\n<td><code>Claude-SearchBot<\/code><\/td>\n<td>Anthropic<\/td>\n<td>Improving search result relevance and accuracy<\/td>\n<td align=\"right\">High<\/td>\n<td>Claude search visibility<\/td>\n<td>Category, docs, comparison, integration, pricing pages<\/td>\n<\/tr>\n<tr>\n<td><code>Claude-User<\/code><\/td>\n<td>Anthropic<\/td>\n<td>User-directed retrieval in Claude<\/td>\n<td align=\"right\">High for live user prompts<\/td>\n<td>Buyer-journey monitoring<\/td>\n<td>User-prompt fetches, status codes, quote-worthy pages<\/td>\n<\/tr>\n<tr>\n<td><code>PerplexityBot<\/code><\/td>\n<td>Perplexity<\/td>\n<td>Surfacing and linking websites in Perplexity results<\/td>\n<td align=\"right\">Very high<\/td>\n<td>Perplexity citation visibility<\/td>\n<td>Verified IPs, pages fetched, citations gained or lost<\/td>\n<\/tr>\n<tr>\n<td><code>Perplexity-User<\/code><\/td>\n<td>Perplexity<\/td>\n<td>User-triggered fetches inside Perplexity<\/td>\n<td align=\"right\">High, with governance risk<\/td>\n<td>WAF and policy review<\/td>\n<td>Fetch paths, IP verification, suspicious undeclared traffic<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>The Most Important Distinction: Training, Search, and User Fetching<\/h2>\n<p>AI crawlers fall into three operational groups. Mixing them together creates bad policy decisions.<\/p>\n<table>\n<thead>\n<tr>\n<th>Bot type<\/th>\n<th>Examples<\/th>\n<th>What it affects<\/th>\n<th>Robots.txt decision<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Training crawler<\/strong><\/td>\n<td><code>GPTBot<\/code>, <code>ClaudeBot<\/code><\/td>\n<td>Whether public content may be collected for model training or model improvement<\/td>\n<td>Decide with legal, privacy, content, and brand teams<\/td>\n<\/tr>\n<tr>\n<td><strong>Search crawler<\/strong><\/td>\n<td><code>OAI-SearchBot<\/code>, <code>Claude-SearchBot<\/code>, <code>PerplexityBot<\/code><\/td>\n<td>Whether an AI search product can discover and use public pages<\/td>\n<td>Usually allow approved public pages if AI visibility matters<\/td>\n<\/tr>\n<tr>\n<td><strong>User-triggered fetcher<\/strong><\/td>\n<td><code>ChatGPT-User<\/code>, <code>Claude-User<\/code>, <code>Perplexity-User<\/code><\/td>\n<td>Whether an AI assistant can fetch a page during a user&#39;s query<\/td>\n<td>Monitor closely; do not assume behavior matches scheduled crawling<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>OpenAI&#39;s crawler documentation states that <code>OAI-SearchBot<\/code> is for ChatGPT search, while <code>GPTBot<\/code> is for content that may be used in training foundation models. OpenAI also says each setting is independent, so a site can allow <code>OAI-SearchBot<\/code> while disallowing <code>GPTBot<\/code>. See <a href=\"https:\/\/developers.openai.com\/api\/docs\/bots\" target=\"_blank\" rel=\"noopener\">OpenAI&#39;s crawler documentation<\/a>.<\/p>\n<p>Anthropic documents a similar split: <code>ClaudeBot<\/code> supports model training, <code>Claude-SearchBot<\/code> improves search result relevance, and <code>Claude-User<\/code> supports user-directed retrieval. See <a href=\"https:\/\/support.claude.com\/en\/articles\/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler\" target=\"_blank\" rel=\"noopener\">Anthropic&#39;s crawler documentation<\/a>.<\/p>\n<p>Perplexity states that <code>PerplexityBot<\/code> is designed to surface and link websites in Perplexity search results and is not used to crawl content for AI foundation models. See <a href=\"https:\/\/docs.perplexity.ai\/docs\/resources\/perplexity-crawlers\" target=\"_blank\" rel=\"noopener\">Perplexity&#39;s crawler documentation<\/a>.<\/p>\n<h2>What GPTBot Does for Brand Visibility<\/h2>\n<p><code>GPTBot<\/code> is not the master switch for ChatGPT visibility. It is OpenAI&#39;s training-related crawler. Disallowing it signals that your content should not be used in training OpenAI generative AI foundation models.<\/p>\n<p>That matters for governance. It matters less for whether your page appears in a ChatGPT search answer this week. OpenAI says <code>OAI-SearchBot<\/code>, not <code>GPTBot<\/code>, is used to surface websites in ChatGPT search features. OpenAI also notes that sites opted out of <code>OAI-SearchBot<\/code> will not be shown in ChatGPT search answers, although they may still appear as navigational links.<\/p>\n<p>Use this policy split:<\/p>\n<table>\n<thead>\n<tr>\n<th>Business goal<\/th>\n<th><code>GPTBot<\/code> policy<\/th>\n<th><code>OAI-SearchBot<\/code> policy<\/th>\n<th><code>ChatGPT-User<\/code> policy<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Maximize ChatGPT search visibility<\/td>\n<td>Allow or selectively allow<\/td>\n<td>Allow approved public pages<\/td>\n<td>Allow public pages and monitor logs<\/td>\n<\/tr>\n<tr>\n<td>Opt out of training use<\/td>\n<td>Disallow all or sensitive paths<\/td>\n<td>Allow public marketing and docs pages<\/td>\n<td>Allow public pages only<\/td>\n<\/tr>\n<tr>\n<td>Protect proprietary content<\/td>\n<td>Disallow sensitive paths and enforce auth<\/td>\n<td>Disallow private paths<\/td>\n<td>Require authentication for private pages<\/td>\n<\/tr>\n<tr>\n<td>Reduce bot load<\/td>\n<td>Rate-limit or disallow low-value paths<\/td>\n<td>Keep high-value pages accessible<\/td>\n<td>Watch bursts and WAF actions<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The common mistake is blocking <code>GPTBot<\/code>, seeing no immediate change in ChatGPT search visibility, and assuming crawler policy does not matter. The better diagnosis is that the wrong crawler was measured.<\/p>\n<h2>What ClaudeBot Does for Brand Visibility<\/h2>\n<p><code>ClaudeBot<\/code> is Anthropic&#39;s training-related crawler. It helps collect public web content that could contribute to model training. For current Claude visibility, however, <code>Claude-SearchBot<\/code> and <code>Claude-User<\/code> often matter more.<\/p>\n<p>Anthropic says disabling <code>Claude-SearchBot<\/code> may reduce visibility and accuracy in user search results. It also says disabling <code>Claude-User<\/code> prevents its system from retrieving your content in response to a user query, which may reduce visibility for user-directed web search.<\/p>\n<p>For a SaaS or technical site, monitor these Claude-related events separately:<\/p>\n<ul>\n<li><code>ClaudeBot<\/code> hitting public pages that legal or content teams do not want used for training.<\/li>\n<li><code>Claude-SearchBot<\/code> reaching category, docs, comparison, integration, and pricing pages.<\/li>\n<li><code>Claude-User<\/code> fetching pages after buyer-style prompts.<\/li>\n<li>403, 406, 429, and WAF challenges on pages Claude needs to read.<\/li>\n<li>Whether Claude later cites the page or describes the brand accurately.<\/li>\n<\/ul>\n<p>If Claude understands your product category incorrectly, crawler access alone will not fix it. Strengthen entity clarity, comparison content, third-party corroboration, and crawlable HTML.<\/p>\n<h2>What PerplexityBot Does for Brand Visibility<\/h2>\n<p><code>PerplexityBot<\/code> is the most directly visibility-oriented of the three named bots. Perplexity says it is used to surface and link websites in Perplexity results. If it cannot fetch your best pages, Perplexity has fewer first-party sources to cite.<\/p>\n<p>That makes <code>PerplexityBot<\/code> a high-priority crawler for brands tracking AI citations. Watch whether it reaches:<\/p>\n<ul>\n<li>Homepage and product pages.<\/li>\n<li>Category and use-case pages.<\/li>\n<li>Comparison and alternatives pages.<\/li>\n<li>Public documentation and integration pages.<\/li>\n<li>Pricing, security, trust, and case study pages where appropriate.<\/li>\n<\/ul>\n<p>The governance issue is verification. Perplexity also documents <code>Perplexity-User<\/code> for user-triggered fetches and says this fetcher generally ignores robots.txt because the user requested the fetch. In August 2025, Cloudflare reported that Perplexity used undeclared crawling behavior after blocks, including a generic Chrome-like user agent, IPs outside declared ranges, and traffic observed at <strong>3-6 million daily stealth requests<\/strong> compared with <strong>20-25 million daily declared <code>Perplexity-User<\/code> requests<\/strong> in Cloudflare&#39;s environment. See Cloudflare&#39;s investigation, <a href=\"https:\/\/blog.cloudflare.com\/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives\/\" target=\"_blank\" rel=\"noopener\">Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives<\/a>.<\/p>\n<p>The practical rule: treat <code>PerplexityBot<\/code> as important for visibility, but verify by <strong>user agent plus official IP range<\/strong>, not by user-agent string alone.<\/p>\n<h2>MaxAEO&#39;s Crawler Visibility Relevance Score<\/h2>\n<p>Crawler Visibility Relevance, or CVR, is a practical scoring model for deciding which AI crawlers deserve monitoring, access, and executive reporting. It ranks bots by their path to measurable AI visibility, not by traffic volume.<\/p>\n<p>Use a 100-point score:<\/p>\n<table>\n<thead>\n<tr>\n<th>Factor<\/th>\n<th align=\"right\">Weight<\/th>\n<th>Question to answer<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Retrieval role<\/td>\n<td align=\"right\">30<\/td>\n<td>Does this bot feed current search or only training?<\/td>\n<\/tr>\n<tr>\n<td>Citation path<\/td>\n<td align=\"right\">20<\/td>\n<td>Can fetched pages become cited sources in visible answers?<\/td>\n<\/tr>\n<tr>\n<td>Log verifiability<\/td>\n<td align=\"right\">20<\/td>\n<td>Are user agents and IP ranges published and testable?<\/td>\n<\/tr>\n<tr>\n<td>Recency value<\/td>\n<td align=\"right\">15<\/td>\n<td>Can the bot discover fresh launches, pricing, docs, and claims?<\/td>\n<\/tr>\n<tr>\n<td>Governance risk<\/td>\n<td align=\"right\">15<\/td>\n<td>Does allowing it create privacy, legal, load, or policy risk?<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Applied to current public documentation, a typical B2B SaaS priority order looks like this:<\/p>\n<table>\n<thead>\n<tr>\n<th>Bot group<\/th>\n<th align=\"right\">CVR score<\/th>\n<th>Why<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>PerplexityBot<\/code><\/td>\n<td align=\"right\">82<\/td>\n<td>Direct path to Perplexity search results and citations, with verification risk<\/td>\n<\/tr>\n<tr>\n<td><code>OAI-SearchBot<\/code><\/td>\n<td align=\"right\">80<\/td>\n<td>Direct path to ChatGPT search visibility and separate from <code>GPTBot<\/code><\/td>\n<\/tr>\n<tr>\n<td><code>Claude-SearchBot<\/code><\/td>\n<td align=\"right\">74<\/td>\n<td>Direct path to Claude search relevance and answer accuracy<\/td>\n<\/tr>\n<tr>\n<td><code>ChatGPT-User<\/code>, <code>Claude-User<\/code>, <code>Perplexity-User<\/code><\/td>\n<td align=\"right\">70-85<\/td>\n<td>High impact when buyers ask live questions, but behavior differs from scheduled crawling<\/td>\n<\/tr>\n<tr>\n<td><code>GPTBot<\/code><\/td>\n<td align=\"right\">45<\/td>\n<td>Important for training permissions, less direct for current search citations<\/td>\n<\/tr>\n<tr>\n<td><code>ClaudeBot<\/code><\/td>\n<td align=\"right\">45<\/td>\n<td>Important for training permissions, separate from Claude search retrieval<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The score should change by business model. A publisher may weight governance risk higher. A developer tool may weight <code>Claude-User<\/code> higher because Claude is heavily used for technical workflows. A cybersecurity company may weight log verifiability and WAF handling higher than citation upside.<\/p>\n<h2>Allow, Block, or Segment: Decision Matrix<\/h2>\n<p>A good AI crawler policy is segmented by page type. Public pages that help buyers understand your brand should usually be accessible to search and user-triggered AI systems. Sensitive, gated, duplicate, or low-value content should not be exposed just because an AI bot asks for it.<\/p>\n<table>\n<thead>\n<tr>\n<th>Page type<\/th>\n<th>Recommended policy<\/th>\n<th>Reason<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Homepage<\/td>\n<td>Allow search bots and user fetchers<\/td>\n<td>Core entity understanding<\/td>\n<\/tr>\n<tr>\n<td>Product pages<\/td>\n<td>Allow approved public pages<\/td>\n<td>Product-category association<\/td>\n<\/tr>\n<tr>\n<td>Category and use-case pages<\/td>\n<td>Allow<\/td>\n<td>Helps AI engines know when to recommend you<\/td>\n<\/tr>\n<tr>\n<td>Comparison pages<\/td>\n<td>Allow<\/td>\n<td>Important for answer-engine shortlists<\/td>\n<\/tr>\n<tr>\n<td>Public docs and help center<\/td>\n<td>Allow if accurate and maintained<\/td>\n<td>Strong citation candidates<\/td>\n<\/tr>\n<tr>\n<td>Pricing page<\/td>\n<td>Case by case<\/td>\n<td>Useful for buyers but compliance-sensitive<\/td>\n<\/tr>\n<tr>\n<td>Security and trust pages<\/td>\n<td>Allow public summaries; protect private docs<\/td>\n<td>Enterprise evaluation signal<\/td>\n<\/tr>\n<tr>\n<td>Case studies<\/td>\n<td>Allow if approved<\/td>\n<td>Proof for AI-generated recommendations<\/td>\n<\/tr>\n<tr>\n<td>Gated assets<\/td>\n<td>Do not expose only for crawlers<\/td>\n<td>Keep access logic consistent<\/td>\n<\/tr>\n<tr>\n<td>Internal search, admin, staging, parameters<\/td>\n<td>Block and enforce<\/td>\n<td>Low visibility value, high crawl waste<\/td>\n<\/tr>\n<tr>\n<td>Confidential files<\/td>\n<td>Block with authentication, not robots.txt alone<\/td>\n<td>Robots.txt is not security<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>For broader policy design, use <a href=\"https:\/\/maxaeo.ai\/blog\/block-or-allow-ai-crawlers\">Block or Allow AI Crawlers? A Per-Bot Decision Guide<\/a> before editing production rules.<\/p>\n<h2>Example robots.txt Patterns<\/h2>\n<p>A selective policy can allow search-oriented AI crawlers while opting out of training-oriented crawlers:<\/p>\n<pre><code class=\"language-txt\">User-agent: GPTBot\nDisallow: \/\n\nUser-agent: ClaudeBot\nDisallow: \/\n\nUser-agent: OAI-SearchBot\nAllow: \/\n\nUser-agent: Claude-SearchBot\nAllow: \/\n\nUser-agent: PerplexityBot\nAllow: \/\n\nUser-agent: *\nDisallow:\n<\/code><\/pre>\n<p>That example is not universal. It fits teams that want AI search visibility while declining some training-related access.<\/p>\n<p>A more conservative policy might allow only selected public directories:<\/p>\n<pre><code class=\"language-txt\">User-agent: OAI-SearchBot\nAllow: \/blog\/\nAllow: \/docs\/\nAllow: \/compare\/\nAllow: \/integrations\/\nDisallow: \/\n\nUser-agent: Claude-SearchBot\nAllow: \/blog\/\nAllow: \/docs\/\nAllow: \/compare\/\nAllow: \/integrations\/\nDisallow: \/\n\nUser-agent: PerplexityBot\nAllow: \/blog\/\nAllow: \/docs\/\nAllow: \/compare\/\nAllow: \/integrations\/\nDisallow: \/\n<\/code><\/pre>\n<p>Before shipping this, test the exact matching behavior. Google&#39;s robots.txt guide notes that robots.txt rules apply only to the <strong>host, protocol, and port<\/strong> where the file is hosted. The formal standard is <a href=\"https:\/\/datatracker.ietf.org\/doc\/html\/rfc9309\" target=\"_blank\" rel=\"noopener\">RFC 9309, the Robots Exclusion Protocol<\/a>, and Google&#39;s interpretation is documented in <a href=\"https:\/\/developers.google.com\/crawling\/docs\/robots-txt\/robots-txt-spec\" target=\"_blank\" rel=\"noopener\">How Google interprets the robots.txt specification<\/a>.<\/p>\n<h2>How to Verify GPTBot, ClaudeBot, and PerplexityBot in Logs<\/h2>\n<p>Do not report AI crawler activity by user agent alone. User-agent strings can be spoofed. A reliable workflow verifies identity, classifies purpose, and ties access to AI answer outcomes.<\/p>\n<p>Use this six-step log process:<\/p>\n<ol>\n<li><strong>Collect the right fields:<\/strong> timestamp, URL, status code, user agent, IP, ASN, country, response size, cache status, WAF action, and referrer where available.<\/li>\n<li><strong>Verify identity:<\/strong> match the user agent to official IP ranges or published verification methods.<\/li>\n<li><strong>Classify purpose:<\/strong> label each event as training crawler, search crawler, user-triggered fetcher, ad validation, unknown, or suspicious.<\/li>\n<li><strong>Group by visibility page:<\/strong> homepage, product, category, comparison, docs, pricing, integrations, case studies, and trust pages.<\/li>\n<li><strong>Compare with answer outcomes:<\/strong> check whether ChatGPT, Claude, Perplexity, Gemini, Copilot, Grok, Google AI Mode, and AI Overviews mention, cite, and rank the brand after crawl events.<\/li>\n<li><strong>Investigate gaps:<\/strong> if access exists but citations do not, look for weak content, client-side rendering, blocked resources, outdated claims, or missing third-party corroboration.<\/li>\n<\/ol>\n<p>Use official verification sources where available:<\/p>\n<table>\n<thead>\n<tr>\n<th>Platform<\/th>\n<th>Verification source<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>OpenAI<\/td>\n<td><code>openai.com\/searchbot.json<\/code>, <code>openai.com\/gptbot.json<\/code>, and <code>openai.com\/chatgpt-user.json<\/code>, linked from the OpenAI crawler docs<\/td>\n<\/tr>\n<tr>\n<td>Anthropic<\/td>\n<td><code>claude.com\/crawling\/bots.json<\/code>, linked from Anthropic&#39;s crawler docs<\/td>\n<\/tr>\n<tr>\n<td>Perplexity<\/td>\n<td><code>perplexity.com\/perplexitybot.json<\/code> and <code>perplexity.com\/perplexity-user.json<\/code>, linked from Perplexity&#39;s crawler docs<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>A minimal log label taxonomy should look like this:<\/p>\n<table>\n<thead>\n<tr>\n<th>Label<\/th>\n<th>Matching examples<\/th>\n<th>Action<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>ai_training<\/code><\/td>\n<td><code>GPTBot<\/code>, <code>ClaudeBot<\/code><\/td>\n<td>Govern by training policy<\/td>\n<\/tr>\n<tr>\n<td><code>ai_search<\/code><\/td>\n<td><code>OAI-SearchBot<\/code>, <code>Claude-SearchBot<\/code>, <code>PerplexityBot<\/code><\/td>\n<td>Keep important public pages accessible<\/td>\n<\/tr>\n<tr>\n<td><code>ai_user_fetch<\/code><\/td>\n<td><code>ChatGPT-User<\/code>, <code>Claude-User<\/code>, <code>Perplexity-User<\/code><\/td>\n<td>Monitor buyer-prompt behavior and WAF results<\/td>\n<\/tr>\n<tr>\n<td><code>unverified_ai_claim<\/code><\/td>\n<td>AI-looking UA without verified IP<\/td>\n<td>Investigate before allowing<\/td>\n<\/tr>\n<tr>\n<td><code>suspicious_fetch<\/code><\/td>\n<td>Browser UA, AI-like access pattern, blocked declared bot<\/td>\n<td>Challenge, rate-limit, or review with security<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Why Bot Access Does Not Guarantee AI Citations<\/h2>\n<p>A crawler hit proves access. It does not prove inclusion, citation, ranking, or recommendation. AI systems may fetch your page and still cite a competitor if your page is thin, stale, vague, hard to parse, or less authoritative.<\/p>\n<p>Strong AI citation candidates usually have:<\/p>\n<ul>\n<li>A clear entity statement: what the company is, who it serves, and which category it belongs to.<\/li>\n<li>Direct answers to buyer questions in short, extractable passages.<\/li>\n<li>Specific product capabilities instead of broad positioning.<\/li>\n<li>Freshness signals for pricing, integrations, benchmarks, and product launches.<\/li>\n<li>Crawlable HTML for critical copy, not client-side-only content.<\/li>\n<li>Consistent claims across product pages, docs, schema, press pages, and third-party profiles.<\/li>\n<li>Comparison content that explains tradeoffs honestly.<\/li>\n<\/ul>\n<p>This is why a brand can rank first in Google and still fail to appear in AI answers. The source-selection mechanics differ. For that gap, see <a href=\"https:\/\/maxaeo.ai\/blog\/rank-google-not-ai-search\">You Rank #1 on Google but Don&#39;t Appear in AI Answers &#8211; Here&#39;s Why<\/a>.<\/p>\n<p>If your key copy is rendered only after JavaScript execution, fix that before blaming crawler policy. The technical risk is covered in <a href=\"https:\/\/maxaeo.ai\/blog\/javascript-ai-search-visibility\">JavaScript AI Search Visibility: Why Client-Side Content Hurts AI Crawlers<\/a>.<\/p>\n<h2>What to Fix Before Changing Crawler Rules<\/h2>\n<p>Crawler access amplifies the quality of what the bot can read. It does not make weak pages cite-worthy.<\/p>\n<table>\n<thead>\n<tr>\n<th>Check<\/th>\n<th>What good looks like<\/th>\n<th>Risk if weak<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Entity clarity<\/td>\n<td>The page states brand, category, audience, and use case plainly<\/td>\n<td>AI engines describe the brand vaguely<\/td>\n<\/tr>\n<tr>\n<td>Crawlable HTML<\/td>\n<td>Main content appears in server-rendered or easily retrievable HTML<\/td>\n<td>Crawlers miss product details<\/td>\n<\/tr>\n<tr>\n<td>Passage quality<\/td>\n<td>Each section answers a specific buyer question<\/td>\n<td>Engines cite clearer competitors<\/td>\n<\/tr>\n<tr>\n<td>Freshness<\/td>\n<td>Dates and updates are visible where recency matters<\/td>\n<td>AI answers use stale positioning<\/td>\n<\/tr>\n<tr>\n<td>Internal consistency<\/td>\n<td>Pricing, product, integrations, and trust claims match across pages<\/td>\n<td>AI descriptions become inconsistent<\/td>\n<\/tr>\n<tr>\n<td>External corroboration<\/td>\n<td>Third-party sources support the same claims<\/td>\n<td>Engines hesitate to recommend the brand<\/td>\n<\/tr>\n<tr>\n<td>Log coverage<\/td>\n<td>Important pages receive verified successful bot hits<\/td>\n<td>Visibility gaps are hard to diagnose<\/td>\n<\/tr>\n<tr>\n<td>Source map<\/td>\n<td>You know which engines rely on which search indexes or sources<\/td>\n<td>Teams over-attribute one crawler to every AI result<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>For multi-engine planning, pair crawler logs with a source-map view. MaxAEO&#39;s guide to <a href=\"https:\/\/maxaeo.ai\/blog\/which-search-engines-power-ai-answers\">which search index powers each AI engine<\/a> explains why one crawler log rarely explains every AI answer surface.<\/p>\n<h2>A Worked Example: From Bot Hits to Visibility Fixes<\/h2>\n<p>A useful AI crawler report connects three datasets: server logs, page inventory, and daily AI answer tracking.<\/p>\n<p>Example method:<\/p>\n<ol>\n<li>Pull 30 days of verified AI crawler and user-fetcher requests.<\/li>\n<li>Filter to commercial pages: homepage, product, category, comparison, docs, pricing, integrations, and case studies.<\/li>\n<li>Mark a request successful only if it returns a 2xx response and meaningful HTML.<\/li>\n<li>Run the same buyer prompts daily across ChatGPT, Claude, Perplexity, Gemini, Copilot, Grok, Google AI Mode, and AI Overviews.<\/li>\n<li>Compare crawl events with AI citations, brand rank, answer sentiment, and description accuracy.<\/li>\n<li>Prioritize fixes where crawl access exists but citations or recommendations are missing.<\/li>\n<\/ol>\n<table>\n<thead>\n<tr>\n<th>Observation<\/th>\n<th>Likely diagnosis<\/th>\n<th>Fix<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>PerplexityBot<\/code> reaches docs pages, but Perplexity cites directories<\/td>\n<td>Docs explain features but not buyer comparisons<\/td>\n<td>Add concise comparison and use-case sections<\/td>\n<\/tr>\n<tr>\n<td><code>OAI-SearchBot<\/code> reaches the homepage, but ChatGPT omits the brand from category shortlists<\/td>\n<td>Weak category association<\/td>\n<td>Strengthen category pages and third-party corroboration<\/td>\n<\/tr>\n<tr>\n<td><code>Claude-SearchBot<\/code> hits pricing, but Claude gives outdated caveats<\/td>\n<td>Pricing page lacks update signals<\/td>\n<td>Add visible update date and stable pricing explanation<\/td>\n<\/tr>\n<tr>\n<td><code>GPTBot<\/code> is blocked, but ChatGPT search visibility is unchanged<\/td>\n<td>Training crawler was not the immediate search blocker<\/td>\n<td>Audit <code>OAI-SearchBot<\/code> and <code>ChatGPT-User<\/code> access<\/td>\n<\/tr>\n<tr>\n<td>Declared <code>PerplexityBot<\/code> is blocked, but unknown Chrome-like traffic rises<\/td>\n<td>Possible undeclared or proxy fetch behavior<\/td>\n<td>Verify IPs, challenge suspicious traffic, review WAF logs<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The operating difference is simple: <strong>logs tell you what reached the server; AI visibility tracking tells you whether market-facing answers changed.<\/strong><\/p>\n<h2>Where llms.txt Fits<\/h2>\n<p><code>llms.txt<\/code> can help AI systems understand preferred summaries, important pages, and machine-readable context. It does not replace robots.txt, authentication, WAF controls, or content quality.<\/p>\n<p>Use <code>llms.txt<\/code> for guidance. Use robots.txt for crawler directives. Use authentication and authorization for private content. Use log verification to confirm what actually happened.<\/p>\n<p>For implementation details, see <a href=\"https:\/\/maxaeo.ai\/blog\/llms-txt-ai-visibility\">llms.txt for AI Visibility: Does It Actually Work, and How to Write One<\/a>.<\/p>\n<h2>Common Mistakes<\/h2>\n<p><strong>Mistake 1: Treating GPTBot as the ChatGPT visibility switch.<\/strong><br \/>\nOpenAI separates <code>GPTBot<\/code> from <code>OAI-SearchBot<\/code> and <code>ChatGPT-User<\/code>. Training permission and search visibility need separate rules.<\/p>\n<p><strong>Mistake 2: Blocking all AI crawlers because one bot is noisy.<\/strong><br \/>\nThat may reduce risk, but it can also remove useful public pages from answer engines buyers use for shortlist research.<\/p>\n<p><strong>Mistake 3: Allowing bots without measuring outcomes.<\/strong><br \/>\nIf <code>PerplexityBot<\/code> crawls 500 pages and Perplexity still cites outdated third-party sources, the problem may be structure, authority, or source selection.<\/p>\n<p><strong>Mistake 4: Ignoring user-triggered fetchers.<\/strong><br \/>\nWhen a buyer asks an AI assistant to compare vendors, user-triggered agents may fetch pages differently from scheduled crawlers.<\/p>\n<p><strong>Mistake 5: Using robots.txt as security.<\/strong><br \/>\nRobots.txt is a public crawling preference file. Private content belongs behind authentication, authorization, contractual controls, and appropriate WAF rules.<\/p>\n<h2>Recommended Policy for B2B SaaS and Tech Brands<\/h2>\n<p>Most B2B SaaS and technical brands should allow search and user-triggered access to approved public pages, govern training crawlers separately, and block sensitive or low-value paths.<\/p>\n<p>A practical default:<\/p>\n<ul>\n<li>Allow <code>OAI-SearchBot<\/code>, <code>Claude-SearchBot<\/code>, and <code>PerplexityBot<\/code> on public marketing, docs, comparison, integration, and category pages.<\/li>\n<li>Monitor <code>ChatGPT-User<\/code>, <code>Claude-User<\/code>, and <code>Perplexity-User<\/code> because they can reflect real buyer prompts.<\/li>\n<li>Decide separately whether <code>GPTBot<\/code> and <code>ClaudeBot<\/code> may access public content for training-related purposes.<\/li>\n<li>Block admin, staging, internal search, parameter traps, duplicate pages, private files, and sensitive commercial content.<\/li>\n<li>Verify user agent plus IP range before trusting crawler labels.<\/li>\n<li>Review crawler policy monthly and after migrations, product launches, pricing changes, WAF updates, and legal policy changes.<\/li>\n<li>Track the outcome that matters: AI mentions, citations, rank position, sentiment, and description accuracy.<\/li>\n<\/ul>\n<p>MaxAEO is built for that last step: connecting AI search visibility outcomes to the technical and content changes that can move them.<\/p>\n<h2>Frequently Asked Questions<\/h2>\n<h3>What is the difference between GPTBot, ClaudeBot, and PerplexityBot?<\/h3>\n<p><code>GPTBot<\/code> is OpenAI&#39;s training-related crawler, <code>ClaudeBot<\/code> is Anthropic&#39;s training-related crawler, and <code>PerplexityBot<\/code> is Perplexity&#39;s crawler for surfacing and linking websites in Perplexity results. They should not share one blanket robots.txt rule.<\/p>\n<h3>Which crawler matters most for AI brand visibility?<\/h3>\n<p>For current answer visibility, prioritize <code>PerplexityBot<\/code>, <code>OAI-SearchBot<\/code>, <code>Claude-SearchBot<\/code>, and user-triggered fetchers such as <code>ChatGPT-User<\/code>, <code>Claude-User<\/code>, and <code>Perplexity-User<\/code>. For training opt-out and content governance, prioritize <code>GPTBot<\/code> and <code>ClaudeBot<\/code>.<\/p>\n<h3>Should I block GPTBot?<\/h3>\n<p>Block <code>GPTBot<\/code> if your policy is to prevent public content from being used for OpenAI foundation-model training. Do not assume blocking <code>GPTBot<\/code> removes your brand from ChatGPT search answers. For ChatGPT search visibility, monitor <code>OAI-SearchBot<\/code> and <code>ChatGPT-User<\/code>.<\/p>\n<h3>Is ClaudeBot the same as Claude-SearchBot?<\/h3>\n<p>No. Anthropic documents <code>ClaudeBot<\/code> for training-related collection and <code>Claude-SearchBot<\/code> for improving search result relevance and accuracy. If your goal is current Claude visibility, <code>Claude-SearchBot<\/code> and <code>Claude-User<\/code> usually deserve more attention than <code>ClaudeBot<\/code> volume alone.<\/p>\n<h3>Does allowing PerplexityBot guarantee Perplexity citations?<\/h3>\n<p>No. Allowing <code>PerplexityBot<\/code> only makes access possible. Perplexity still chooses sources based on relevance, freshness, retrievability, and answer fit. Improve citations by making key pages crawlable, specific, current, and easy to quote.<\/p>\n<h3>Can llms.txt replace robots.txt?<\/h3>\n<p>No. <code>llms.txt<\/code> is a guidance layer, not an access-control system. Use robots.txt for crawler directives, authentication for private content, WAF rules for enforcement, and <code>llms.txt<\/code> for machine-readable context.<\/p>\n<h3>How often should brand teams review AI crawler logs?<\/h3>\n<p>Review high-level AI crawler trends weekly and run a deeper policy review monthly. Also review logs after site migrations, product launches, pricing changes, documentation restructures, WAF changes, and sudden drops in AI share of voice.<\/p>\n<h2>Final Takeaway<\/h2>\n<p><code>GPTBot<\/code>, <code>ClaudeBot<\/code>, and <code>PerplexityBot<\/code> answer different business questions. <code>GPTBot<\/code> and <code>ClaudeBot<\/code> are mainly training-governance controls. <code>PerplexityBot<\/code>, <code>OAI-SearchBot<\/code>, <code>Claude-SearchBot<\/code>, and user-triggered fetchers are closer to current AI citations and brand visibility.<\/p>\n<p>Do not decide with a blanket allow or block. Segment by bot purpose, page type, verification confidence, and risk. Then measure whether AI engines actually mention, cite, rank, and describe your brand accurately when buyers ask.<\/p>\n<p><script type=\"application\/ld+json\">\n{\n  \"@context\": \"https:\/\/schema.org\",\n  \"@graph\": [\n    {\n      \"@type\": \"Article\",\n      \"headline\": \"GPTBot ClaudeBot PerplexityBot Compared: Which AI Crawlers Matter for Brand Visibility?\",\n      \"description\": \"Compare GPTBot, ClaudeBot, and PerplexityBot by purpose, robots.txt behavior, log verification, and impact on AI search visibility.\",\n      \"author\": {\n        \"@type\": \"Organization\",\n        \"name\": \"maxaeo\"\n      },\n      \"publisher\": {\n        \"@type\": \"Organization\",\n        \"name\": \"maxaeo\"\n      },\n      \"datePublished\": \"2026-07-09\",\n      \"dateModified\": \"2026-07-09\",\n      \"image\": \"image-placeholder\",\n      \"mainEntityOfPage\": {\n        \"@type\": \"WebPage\",\n        \"@id\": \"https:\/\/maxaeo.ai\/blog\/gptbot-claudebot-perplexitybot\"\n      }\n    },\n    {\n      \"@type\": \"FAQPage\",\n      \"mainEntity\": [\n        {\n          \"@type\": \"Question\",\n          \"name\": \"What is the difference between GPTBot, ClaudeBot, and PerplexityBot?\",\n          \"acceptedAnswer\": {\n            \"@type\": \"Answer\",\n            \"text\": \"GPTBot is OpenAI's training-related crawler, ClaudeBot is Anthropic's training-related crawler, and PerplexityBot is Perplexity's crawler for surfacing and linking websites in Perplexity results.\"\n          }\n        },\n        {\n          \"@type\": \"Question\",\n          \"name\": \"Which crawler matters most for AI brand visibility?\",\n          \"acceptedAnswer\": {\n            \"@type\": \"Answer\",\n            \"text\": \"For current answer visibility, prioritize PerplexityBot, OAI-SearchBot, Claude-SearchBot, and user-triggered fetchers such as ChatGPT-User, Claude-User, and Perplexity-User.\"\n          }\n        },\n        {\n          \"@type\": \"Question\",\n          \"name\": \"Should I block GPTBot?\",\n          \"acceptedAnswer\": {\n            \"@type\": \"Answer\",\n            \"text\": \"Block GPTBot if your policy is to prevent public content from being used for OpenAI foundation-model training. For ChatGPT search visibility, monitor OAI-SearchBot and ChatGPT-User separately.\"\n          }\n        },\n        {\n          \"@type\": \"Question\",\n          \"name\": \"Is ClaudeBot the same as Claude-SearchBot?\",\n          \"acceptedAnswer\": {\n            \"@type\": \"Answer\",\n            \"text\": \"No. ClaudeBot is training-related, while Claude-SearchBot improves search result relevance and accuracy. Claude-SearchBot and Claude-User are usually more relevant for current Claude visibility.\"\n          }\n        },\n        {\n          \"@type\": \"Question\",\n          \"name\": \"Does allowing PerplexityBot guarantee Perplexity citations?\",\n          \"acceptedAnswer\": {\n            \"@type\": \"Answer\",\n            \"text\": \"No. Allowing PerplexityBot makes access possible, but Perplexity still chooses sources based on relevance, freshness, retrievability, and answer fit.\"\n          }\n        },\n        {\n          \"@type\": \"Question\",\n          \"name\": \"Can llms.txt replace robots.txt?\",\n          \"acceptedAnswer\": {\n            \"@type\": \"Answer\",\n            \"text\": \"No. llms.txt is a guidance layer, not an access-control system. Use robots.txt for crawler directives and authentication or WAF controls for private or sensitive content.\"\n          }\n        }\n      ]\n    }\n  ]\n}\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Compare GPTBot, ClaudeBot, and PerplexityBot by purpose, robots.txt behavior, log verification, and impact on AI search visibility.<\/p>\n","protected":false},"author":1,"featured_media":1146,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1147","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts\/1147","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/comments?post=1147"}],"version-history":[{"count":0,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/posts\/1147\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/media\/1146"}],"wp:attachment":[{"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/media?parent=1147"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/categories?post=1147"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/maxaeo.ai\/blog\/wp-json\/wp\/v2\/tags?post=1147"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}