All checks

Crawl Rules

The crawl rules the site publishes for search engine bots

What it is

Robots.txt sits at the root of a domain and implements the Robots Exclusion Protocol, telling crawlers which paths to leave alone. It stops bots from hammering a site, but it will not keep a page out of search results, which is what the noindex tag is for.

Why it matters

Because it is a list of things the owner would rather robots did not touch, it occasionally names directories nothing links to: admin panels, staging paths, old exports. It also shows which crawlers the site treats differently from the rest.

Related checks