- Category
- AI search
- Level
- Beginner
- Controlled by
- robots.txt rules per user agent
- Three jobs
- Training, search indexing, user-requested fetches
- Check it free
- Robots.txt Checker
What is an AI crawler?
Like Googlebot, AI crawlers request pages and read their content. What differs is what the content is for. Most AI companies run separate crawlers for separate jobs:
- Training crawlers (such as GPTBot and ClaudeBot) collect content that may be used to train future models.
- Search crawlers (such as OAI-SearchBot, Claude-SearchBot and PerplexityBot) build indexes that AI answers search at question time.
- User agents (such as ChatGPT-User, Claude-User and Perplexity-User) fetch a page because a user asked the assistant to open it.
Why it matters
Blocking a search crawler removes you from that assistant's answers. Blocking a training crawler keeps your content out of future training data but doesn't stop citations. Many sites block AI crawlers without meaning to, through a CDN or firewall bot setting.
How to manage AI crawlers
- Decide your policy per crawler type: most brands allow search crawlers, and choose separately about training.
- Write matching rules in robots.txt.
- Check that your firewall or CDN doesn't block or challenge the crawlers you allow.
- Make sure key content is in the HTML, because most AI crawlers don't run JavaScript.