- Category
- Technical SEO
- Level
- Beginner
- Discovery through
- Links, sitemaps and submissions
- Controlled by
- robots.txt and server responses
- Check it free
- Sitemap Generator
What is crawling?
Crawling is the first step of search. Crawlers start from known URLs, fetch each page, and add newly found links to a queue. They find pages through internal and external links, XML sitemaps, and notifications such as IndexNow (for engines that support it).
What can stop crawling
- robots.txt rules that disallow a path.
- Server errors (5xx), timeouts and very slow responses.
- Links that only exist in JavaScript events or aren't real
<a href>links. - Pages with no links pointing to them (orphan pages).
Best practices
- Link to important pages from other relevant pages.
- Keep an up-to-date XML sitemap and reference it in robots.txt.
- Return correct status codes and fix server errors.
- Watch the Crawl stats and Pages reports in Search Console.