AI search visibility

Which AI bots should I allow?

Decide by URL group: public discovery, model use, or private access.

Published
Reading time
7 min read
Author
Umer Farooq

Short answer

What matters most

Treat public search access and model-use preferences as separate choices. Allow the search crawlers you want to discover public pages; set each publisher’s training or grounding controls according to your content policy; treat user-triggered fetches separately. Apply rules to intentional URL groups, then test robots.txt and the edge response. Use authentication and permissions for restricted content because robots.txt is a request to crawlers, not a security boundary. None of these settings guarantees a citation.

Group pages before choosing crawler rules

A company may publish product guides for anyone to read while keeping customer contracts and account exports behind sign-in. A blanket crawler rule can quietly hide useful public pages, while a disallow line can create a false sense that a private-looking URL is protected.

Start with the page group and the result you want. Public discovery, future model use, user-requested fetching, and confidential access are separate decisions. The map shows which control belongs to each content group before a technical owner edits a rule.

Use the reading guide to jump to the publisher controls, decision matrix, worked example, security boundary, reusable policy brief, or next decision.

A policy map for public pages, licensed material, and private customer data. Public pages get separate search and model-use decisions; licensed material starts with rights checks; private data needs authentication and server-side permissions. All changes are verified across robots.txt, the CDN or firewall, application access, and request logs.A policy map for public pages, licensed material, and private customer data. Public pages get separate search and model-use decisions; licensed material starts with rights checks; private data needs authentication and server-side permissions. All changes are verified across robots.txt, the CDN or firewall, application access, and request logs.
Choose controls by URL group: discovery and model-use preferences are separate, while private data needs real access control.

Separate search, model use, and user requests

The label AI crawler does not identify one universal purpose. Some bots help a service discover public pages for search. Other tokens express a publisher’s preference about future model training or grounding. A user-requested fetch may follow a different path again.

OpenAI documents OAI-SearchBot for ChatGPT search separately from GPTBot, which may collect content for model training. The settings are independent, so a publisher can allow search while disallowing GPTBot. This is a documented control, not a promise that a page will be cited.

Google uses Googlebot controls for access to Search, including its AI features. Google-Extended is a separate robots.txt token for certain training and grounding uses; Google says it does not change Search inclusion or ranking. Do not block Google-Extended if your actual decision concerns Google Search access.

Anthropic and Perplexity also document separate search and user-requested agents. Their guidance shows why you should inspect each provider’s current policy instead of copying one generic list. Perplexity says its Perplexity-User fetcher is initiated by a user and generally ignores robots.txt rules.

OpenAI separates search from training

OAI-SearchBot supports search visibility in ChatGPT, while GPTBot concerns content that may be used for training. OpenAI says the settings work independently and that blocking its search bot affects appearance in search answers.

OpenAI: Overview of OpenAI Crawlers (checked October 9, 2026)

Google Search uses a different control

Google says Googlebot is the access control for Search and its AI features. Google-Extended manages certain training and grounding uses and does not affect Search inclusion or ranking.

Google Search Central: AI features and your website (updated December 10, 2025)

Anthropic documents three bot purposes

ClaudeBot relates to possible model training, Claude-SearchBot to search, and Claude-User to user-directed retrieval. Anthropic says its bots honor robots.txt, so check the provider guidance for the particular agent you mean to control.

Anthropic Help Center: Does Anthropic crawl data from the web? (April 7, 2026)

Use a policy matrix for each URL group

Use the table to record the desired outcome before writing robots.txt rules. The examples are current publisher tokens checked on October 9, 2026. The names do not imply identical behavior across providers.

On a small screen, scroll the table to read every column.

A qualitative crawler policy matrix. Decide by purpose and page group, then verify each provider’s documented behavior.
DecisionExamples and actionBoundary to check
Public search discoveryGooglebot, OAI-SearchBot, Claude-SearchBot, and PerplexityBot: allow only where public discovery is intended.Crawling eligibility does not guarantee indexing, a search result, or an AI citation.
Model training or groundingGPTBot, ClaudeBot, and Google-Extended: decide separately using your content rights and policy.Google-Extended has a different scope from Google Search; recheck each provider’s current definition.
User-requested page fetchChatGPT-User, Claude-User, and Perplexity-User: decide whether user-directed access is acceptable.Robots behavior differs. Some user-initiated fetches may not follow the same crawler rules.
Private or customer contentDo not publish the URL without access checks. Require authentication and server-side authorization.robots.txt is not an access-control mechanism and cannot keep a public URL confidential.

Apply the rule to a hypothetical publisher

Consider a hypothetical software company with public setup guides, licensed partner research, and a private customer portal. Its leaders want the setup guides to remain discoverable, but they have not granted every model-use right for partner research. Customer exports must stay available only to the signed-in account that owns them. This example is a policy exercise, not a client result.

The team can allow Googlebot and selected AI search agents on the public guides, then choose training and grounding preferences per provider and URL group. It should check partner terms before exposing research publicly or permitting model use. Customer exports stay behind authentication and an authorization check at the application or origin. A robots.txt disallow line may reduce compliant crawling, but it cannot replace those checks.

Before applying the policy, the site owner should compare the intended rules with the effective robots.txt on each hostname and the CDN or firewall behavior. If a site-wide default or edge rule overrides the intended exception, the file can look right while the actual request still fails.

Robots.txt does not protect private URLs

The Robots Exclusion Protocol asks compatible crawlers to follow published rules. The standard explicitly says those rules are not a form of access authorization. A URL that returns customer data without checking the requester remains exposed even when robots.txt says not to crawl it.

For restricted content, use sign-in, server-side permissions, and the protections already available at the origin or CDN. For public pages, test the exact URL paths and the response delivered to each intended crawler. Then inspect request logs or the provider’s webmaster tools to confirm that the rule reached the layer doing the blocking or allowing.

Keep visibility expectations modest. Google says a page that meets its technical requirements still is not guaranteed to be crawled, indexed, or served. A crawler allow rule enables a path; it does not promise inclusion or citations.

Crawler rules are not authorization

RFC 9309 defines robots.txt rules as requests to crawlers and explicitly states that they are not a form of access authorization. Use application permissions to protect nonpublic information.

IETF: RFC 9309, Robots Exclusion Protocol (September 2022)

Write a crawler policy brief

Make the decision reviewable by the content owner, site operator, and anyone responsible for licensing or privacy. Keep the brief next to the robots.txt or edge-rule change so a future edit does not erase the reason for it.

  • URL group: list the public, licensed, and restricted routes covered by the decision.
  • Search goal: identify which public pages should be discoverable and which publisher search agents are in scope.
  • Model-use preference: record the allowed or disallowed training and grounding uses by publisher and content owner.
  • User fetch: note whether user-directed page requests are acceptable and how each provider treats its agent.
  • Protection and verification: name the owner for robots.txt, CDN or WAF rules, application access checks, request logs, and rollback.

Choose the next policy decision

Start with one URL group and test the live response before expanding the policy. If ownership, publisher behavior, or edge rules are unclear, I can help review the current setup and define the smallest safe change.

Does llms.txt help AI search?A practical technical SEO checklist for an AI engineering site

Sources

Direct to Umer

Let's talk about your project.

Tell me what you're trying to solve. I'll reply directly by email.