Technical Documentation & Standards

Overview of OpenAI Crawlers & AI Search Indexing

OpenAI uses web crawlers and user agents to perform actions for its products. Learn how OAI-SearchBot and GPTBot work, how to manage your robots.txt, and how to maximize visibility in ChatGPT Search.

ChatGPT Search Visibility

OAI-SearchBot surfaces websites directly inside ChatGPT Search answers. Sites that block this bot will not appear in ChatGPT search results.

Independent Control

Each user agent is independent. You can allow OAI-SearchBot for search inclusion while configuring GPTBot separately.

~24h Propagation Time

When updating your robots.txt file, it typically takes approximately 24 hours for OpenAI systems to refresh and adjust crawl behavior.

OpenAI User Agents & Specifications

Explore the four primary crawlers deployed by OpenAI, their operational purpose, user agent strings, and IP address endpoints.

OAI-SearchBot

Search Discovery

ChatGPT Search Indexing & Surfacing

OAI-SearchBot is for search. OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features. Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links. To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site’s robots.txt file and allowing requests from our published IP ranges.

Used for Foundation Model Training?

No (Used solely for surfacing search answers in ChatGPT)

Robots.txt Configuration Example
User-agent: OAI-SearchBot
Allow: /
Full User-Agent String:
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
Robots.txt Fetch Marker String:
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; robots.txt; +https://openai.com/searchbot

GPTBot

Model Training

Generative AI Foundation Model Training

GPTBot is used to make our generative AI foundation models more useful and safe. It is used to crawl content that may be used in training our generative AI foundation models. Disallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models. When fetching robots.txt files, a robots.txt marker is added to help site owners distinguish those requests.

Used for Foundation Model Training?

Yes (May be used for training foundation models unless disallowed)

Robots.txt Configuration Example
User-agent: GPTBot
Allow: /  # Or Disallow: / to opt-out of training
Full User-Agent String:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Robots.txt Fetch Marker String:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; robots.txt; +https://openai.com/gptbot

OAI-AdsBot

Ad Verification

ChatGPT Ad Validation & Landing Page Quality

OAI-AdsBot is used to validate the safety of web pages submitted as ads on ChatGPT. When you submit an ad, OpenAI may visit the landing page to ensure it complies with policies. Content from the landing page may also be used to determine when it’s most relevant to show the ad to users. OAI-AdsBot only visits pages submitted as ads, and data collected is not used to train foundation models.

Used for Foundation Model Training?

No (Data collected is never used to train generative models)

Robots.txt Configuration Example
User-agent: OAI-AdsBot
Allow: /
Full User-Agent String:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot

ChatGPT-User

User Invoked

User-Triggered Live Web Browsing & Custom GPT Actions

OpenAI also uses ChatGPT-User for certain user actions in ChatGPT and Custom GPTs. When users ask ChatGPT or a CustomGPT a question, it may visit a web page with a ChatGPT-User agent. ChatGPT users may also interact with external applications via GPT Actions. ChatGPT-User is not used for crawling the web in an automatic fashion. Because these actions are initiated by a user, robots.txt rules may not apply.

Used for Foundation Model Training?

No (Direct browsing requests initiated in real time by end users)

Robots.txt Configuration Example
Note: User-initiated browsing actions bypass automatic crawl rules. Use OAI-SearchBot for search opt-outs.
Full User-Agent String:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
Best Practice Configuration

Optimizing Your Robots.txt for ChatGPT & Google Search

Official OpenAI Documentation

To ensure your website, brand profiles, and marketing content are indexed and cited by both Google and ChatGPT search queries, you can include the following rules in your site’s robots.txt:

# Allow ChatGPT Search to index and cite public pages

User-agent: OAI-SearchBot

Allow: /

# Allow Google Search and major search engines

User-agent: Googlebot

Allow: /

# Manage Generative AI training crawler

User-agent: GPTBot

Allow: /

# Global default

User-agent: *

Allow: /

Disallow: /admin/

Disallow: /api/

Subscribe to OpenAI’s official crawler updates:RSS Feed
Explore GoSocio Plans