Technical SEO · Beginner

robots.txt

robots.txt is a plain-text file at the root of a website that tells crawlers which parts of the site they may or may not visit. It can allow or block AI crawlers individually.

Level
Beginner
Location
yourdomain.com/robots.txt
Standard
RFC 9309 (2022)
Google reads
The first 500 KiB
Check it free
Robots.txt Checker

What is robots.txt?

robots.txt is a set of rules grouped by crawler. Each group starts with User-agent: and lists Allow: and Disallow: paths. Crawlers read the group for their own name, or the * group if none matches, and follow the most specific rule for each URL. The format became an internet standard, RFC 9309, in 2022.

User-agent: *
Disallow: /admin/

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

What it does and doesn't do

  • It controls crawling, not indexing. A blocked page can still be indexed if other sites link to it; use noindex to keep a page out of results.
  • It's a request, not security. Well-behaved crawlers follow it; others may not.
  • It can point crawlers to your sitemaps.

Common mistakes

  • Disallow: / left over from a staging site, blocking everything.
  • Blocking CSS or JavaScript that pages need to render.
  • Using noindex inside robots.txt, which Google stopped supporting in 2019.
  • Blocking AI answer crawlers without meaning to.

Frequently asked questions

What happens if robots.txt returns an error?

If it returns a server error, Google may treat the whole site as disallowed until it's fixed. A 404 means there are no restrictions.

Browse the glossary

Chat on WhatsApp