Blog AI visibility
How ChatGPT and Perplexity decide which sites to cite
November 10, 2026
When someone asks ChatGPT or Perplexity a question, the sources listed under the answer were not chosen by a ranking page the way Google's results are. They were pulled in by a separate crawler, filtered for relevance, and lifted for specific sentences a model judged quotable. Understanding that process is the difference between hoping you show up and actually engineering for it.
Three bots, three jobs
OpenAI runs three separate crawlers, and each has a narrow purpose. Its own documentation puts it plainly: "OAI-SearchBot is used to surface websites in search results in ChatGPT's search features." That is a different bot from GPTBot, which trains the underlying model, and from ChatGPT-User, which fetches a page live when a person asks the assistant to browse it. You can allow one and block another in robots.txt; the settings are independent.
Perplexity works the same way. Its documentation describes PerplexityBot as follows: "PerplexityBot is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." If that bot cannot reach your pages, you are not in the running at all, regardless of how good the content is.
What actually gets lifted into an answer
Being crawlable only gets you a seat at the table. A 2024 Princeton-led study on generative engine optimization (arXiv:2311.09735, presented at KDD) tested which content changes moved the needle most. The single most effective one was adding citations and direct quotations from credible sources: it boosted a page's visibility in generative answers by up to 40%. Plain, well-sourced statements beat clever marketing copy almost every time.
Scale: why this is not a niche channel
This is no longer a rounding error in your traffic. OpenAI said at its October 2025 DevDay that ChatGPT had passed 800 million weekly active users. A meaningful share of your prospects are asking an assistant before they ever open a search engine.
The practical checklist
- Allow GPTBot, OAI-SearchBot, ChatGPT-User, and PerplexityBot in robots.txt (check your security plugin or CDN is not blocking them by default).
- Write at least one direct, factual, number-backed paragraph per key page: a stat, a named source, a plain answer.
- Keep pages server-rendered and reachable without JavaScript, since these bots behave more like classic crawlers than browsers.
- State who you are and what you offer in plain text near the top of the page, not only in a hero image.
Quick answers
Does ChatGPT use the same crawler for training and for citing sources?
No. GPTBot collects training data, OAI-SearchBot powers ChatGPT's search features, and ChatGPT-User fetches a page live when a user asks the assistant to open a link. Each can be allowed or blocked separately in robots.txt.
Will blocking GPTBot stop my site appearing in ChatGPT's answers?
Not necessarily. Blocking GPTBot opts you out of training data, but OAI-SearchBot is the crawler that actually surfaces your pages as sources, so it needs its own, separate allow rule.
Run a free scan at RankMerlin to see whether your site currently blocks any of these crawlers, no account needed. It checks SEO and AI visibility side by side so you know exactly what to fix first.
Newsletter
Enjoyed this article?
Subscribe and new SEO and AI-visibility guides land straight in your inbox. No spam, unsubscribe anytime.