Technical SEO · Beginner

Crawler

A crawler is software that visits web pages to read and index them. Search engines and AI companies run their own, such as Googlebot, Bingbot, OAI-SearchBot, GPTBot, PerplexityBot and ClaudeBot. A site's robots.txt tells them what they may access.

Level
Beginner
Also called
Bot, spider
Examples
Googlebot, Bingbot, OAI-SearchBot, GPTBot, ClaudeBot
Identified by
Its user agent string
Check it free
Robots.txt Checker

What is a crawler?

A crawler (also called a bot or spider) is an automated program that requests web pages, reads their content and follows their links to find more pages. Search engines use crawlers to build the index they rank results from; AI companies use them to build search indexes for AI answers and to collect training data.

Each crawler identifies itself with a user agent name, which is what robots.txt rules refer to.

Why it matters

If a crawler can't reach a page, the system it feeds can't use it. Crawl problems, such as blocked paths, server errors, slow responses or content hidden behind JavaScript, quietly limit what search engines and AI assistants know about your site.

Best practices

  1. Keep robots.txt rules deliberate and tested.
  2. Make sure your firewall or CDN doesn't block the crawlers you want.
  3. Serve important content in the HTML.
  4. Link every page you want found from other pages on your site.

Frequently asked questions

Can someone pretend to be Googlebot?

Yes, user agents can be faked. Google and Bing publish ways to verify their real crawlers, such as reverse DNS lookups and IP lists.

Browse the glossary

Chat on WhatsApp