Skip to content

CHATGPT

ChatGPT citation reference —
three crawlers you must
tell apart

OpenAI gives its training, search, and user-request crawlers different names. Which of the three you allow or block decides whether you can appear in a ChatGPT answer at all.

SHORT ANSWER

ChatGPT is an answer engine operated by OpenAI. The first gate on whether you get cited is crawler access, and OpenAI splits its crawlers into three by purpose — GPTBot collects training data, OAI-SearchBot builds the search index, and ChatGPT-User fetches a page on the spot when a user asks for it. If appearing in search-backed answers is the goal, the decisive one is OAI-SearchBot.

Three distinct crawlersInline source linksNo guaranteed placement (OpenAI states this)

SPEC

ChatGPT specifications

Only items that can be verified are listed. Rows with a verification method on the right can be reproduced yourself.

Operator
OpenAI
Crawlers
GPTBot (training data) · OAI-SearchBot (search index) · ChatGPT-User (live fetch on user request) Check the User-Agent string in your server access log.
Gate for search visibility
OAI-SearchBot. Block this one and you disappear from search-backed answers Open yourdomain.com/robots.txt and look for a Disallow that applies to OAI-SearchBot.
Blocking training while allowing search
Possible. Because the names are separate you can Disallow GPTBot and still Allow OAI-SearchBot
Index sources
OpenAI uses the index built by OAI-SearchBot together with third-party search providers (its official documentation names Bing and Shopify)
How citations appear
Source links are attached inline in the answer text. Count this separately from cases where the brand name appears with no link
Korean
Supported. Even for the same intent, the answer and its sources shift with the exact Korean wording, so we repeat measurements across phrasings
Any guarantee?
OpenAI states in its official documentation that there is no way to guarantee top placement

CONDITIONS

What it takes to be cited in ChatGPT

These are the items that differ most in this engine. Principles common to every engine are collected on the answer engines hub.

01

OAI-SearchBot has to be able to read the page

If robots.txt blocks this crawler, the page drops out of the candidate pool for search-backed answers regardless of content quality. It is the first item we check in an AEO audit.

02

The body text has to exist in the HTML source

Body text painted in by JavaScript can look like an empty page to a crawler. What a human sees and what a crawler reads have to be verified separately.

03

A paragraph has to answer one question completely

Answer engines do not use a document whole; they cut pieces out of it. A cut paragraph has to still make sense on its own to work as a citable unit.

04

The same fact should be verifiable in more than one place

A claim that exists only on your own page is weaker evidence than a fact that is confirmed repeatedly across different sources.

MYTHS

Common misconceptions

Only the ones we meet repeatedly in audits.

  • “Allowing GPTBot is enough to show up in ChatGPT search”

    GPTBot is for training. The route into search-backed answers is the index built by OAI-SearchBot. Confusing the two and opening only the training crawler is a genuinely common configuration.

  • “Publishing more articles raises the odds of being cited”

    Several documents answering the same question compete with each other for the same slot. What has to grow is not the number of documents but the number of distinct questions they answer.

  • “If you show up in ChatGPT you will show up in Google”

    The indexes are different. Google ranking and ChatGPT citation run through separate pipelines, and it is common for one to move while the other stays flat.

HOW TO CHECK

How to check it yourself

A Navirang audit follows the same order. There is nothing stopping you from running it internally first.

  1. 01

    Open yourdomain.com/robots.txt in a browser and check the rules that apply to OAI-SearchBot, GPTBot and ChatGPT-User.

  2. 02

    In a logged-out session, ask a question your brand should be the answer to, and record brand mentions in the answer text and your own domain in the source list **separately**.

  3. 03

    Repeat the same question with different phrasings. Never judge from a single response.

  4. 04

    Check your server log for OAI-SearchBot visits and the timestamp of the most recent one.

FAQ

Questions about ChatGPT

Each answer is written to be quoted as it stands.

Can you get our brand to appear in ChatGPT?

We can build the conditions for it; we cannot guarantee it. OpenAI itself states in its official documentation that there is no way to guarantee top placement. Three things are actually within reach — open the site so OAI-SearchBot can read it, write the documents that should be the answer in citable units, and measure the same questions under the same conditions repeatedly so changes are recorded. If a proposal promises a guarantee, it is worth asking what that promise rests on.

If we block GPTBot, do we disappear from ChatGPT search too?

No. GPTBot collects training data, while the route into search-backed answers is the index built by OAI-SearchBot. For a publisher whose content is itself the product, blocking only GPTBot while allowing OAI-SearchBot is a valid choice. The reverse is not: blocking OAI-SearchBot removes you from search-backed answers regardless of training.

Once we open robots.txt, how soon are we cited?

Not immediately. The crawler has to come back, read the document, and have it reflected in the index, and the revisit interval differs by site. Submitting a sitemap and using IndexNow can bring the revisit forward. Navirang records a baseline at the start of an engagement and reports what changed and when, with the raw observations attached.

What is the difference between a mention and a citation?

A mention is your brand name appearing in the answer text; a citation is your domain appearing in the source list. Because ChatGPT attaches sources inline, the two can be told apart by eye. They move independently — a state where the name is known but the document is not used as evidence, and a state where the document is cited but the brand is not foregrounded, call for different work.

RELATED

Related reading

The canonical crawler list is kept as a table in a separate article — this page carries only the rows for this engine.

Last verified August 27, 2026

How does your brand look in ChatGPT right now?

Send us a URL and our free audit asks the questions in all 7 engines and reports what actually came back, recorded engine by engine.

We reply within one business day.

Free audit Call Email Blog