- Category
- Technical SEO
- Level
- Beginner
- Location
- yourdomain.com/robots.txt
- Standard
- RFC 9309 (2022)
- Google reads
- The first 500 KiB
- Check it free
- Robots.txt Checker
What is robots.txt?
robots.txt is a set of rules grouped by crawler. Each group starts with User-agent: and lists Allow: and Disallow: paths. Crawlers read the group for their own name, or the * group if none matches, and follow the most specific rule for each URL. The format became an internet standard, RFC 9309, in 2022.
User-agent: *
Disallow: /admin/
User-agent: GPTBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
What it does and doesn't do
- It controls crawling, not indexing. A blocked page can still be indexed if other sites link to it; use noindex to keep a page out of results.
- It's a request, not security. Well-behaved crawlers follow it; others may not.
- It can point crawlers to your sitemaps.
Common mistakes
Disallow: /left over from a staging site, blocking everything.- Blocking CSS or JavaScript that pages need to render.
- Using
noindexinside robots.txt, which Google stopped supporting in 2019. - Blocking AI answer crawlers without meaning to.