Definition:
Robots.txt is a text file located at a site’s root that tells crawlers which paths they may request and which they should not. Its rules form part of the Robots Exclusion Protocol and apply to crawlers that recognise them.
The file controls automated access for crawling. It is not linked from the HTML, does not protect private content and does not guarantee that a URL will disappear from search results.
Table of contents
How robots.txt works
A crawler requests the file before fetching other URLs from the site:
- Request: the bot attempts to download
/robots.txtfrom the host it intends to crawl. - Identification: it compares its identifier with the groups defined through
User-agent. - Rule selection: it applies the
AllowandDisallowdirectives in the most specific group that matches it. - Path matching: it compares each URL with the declared paths and uses the most specific match.
- Crawl decision: it requests or avoids the URL according to the result and the behaviour implemented by that crawler.
The file is public and its instructions are not an authorisation mechanism. A malicious bot can ignore them, and anyone can inspect the paths listed in it.
Location and scope
The REP standard specifies that the file must be named robots.txt, use lowercase letters and be available at the top-level path of the service.
- Site root: for
https://www.example.com/, the location ishttps://www.example.com/.robots.txt - Host: rules on
www.example.comdo not automatically apply toshop.example.com. - Protocol: a file served over HTTPS does not by itself define the rules for the HTTP version.
- Port: a non-standard port has its own scope.
- Format: it must be served as UTF-8 plain text; a file placed inside a folder does not control the whole site.
The path is case-sensitive. /robots.txt is the expected location, whereas /Robots.txt may be treated as a different resource.
Directives and syntax
Lines are arranged into groups containing one or more agents and their rules:
User-agent: identifies the crawler addressed by the group. The value*represents agents without a more specific group.Disallow: identifies a path that the agent should not request. An empty value does not block any path.Allow: permits a specific path within a broader blocked path.Sitemap: declares the absolute URL of a sitemap. This line does not belong to an agent group.- Comments: the
#character causes the rest of the line to be interpreted as a comment.
Allow and Disallow paths start with /. Specificity determines which rule applies when several rules match. Google’s robots.txt syntax documentation also supports * for a sequence of characters and $ to mark the end of a URL.
Robots.txt examples
Allow all crawling
User-agent: *
Disallow:
Leaving the path after Disallow empty imposes no restriction.
Block a directory
User-agent: *
Disallow: /private/
The rule asks compatible agents not to crawl URLs whose paths begin with /private/.
Allow an exception inside a blocked directory
User-agent: *
Disallow: /resources/
Allow: /resources/public/
The more specific match allows /resources/public/ to be crawled even though its parent directory is blocked.
Block the entire site and declare a sitemap
User-agent: *
Disallow: /
Sitemap: https://www.example.com/sitemap.xml
Disallow: / requests that all crawling be blocked for that group. The Sitemap line can still be present, but it does not override the block.
Crawling, indexing and security
Robots.txt should be distinguished from other controls:
- Crawling: the rules govern whether a compatible agent may request the content at a path.
- Indexing: a blocked URL may still appear in results if a search engine discovers it through links or other sources.
- Noindex: this directive is placed in a meta tag or HTTP header to request exclusion from the index. The search engine needs to crawl the URL to see it.
- Privacy: a path listed in robots.txt remains public and directly accessible. Confidential content requires authentication or access controls.
- Removal: blocking crawling does not erase a known URL or remove copies stored by other systems.
Google’s robots.txt guide explicitly states that the file manages crawl traffic and is not a mechanism for keeping a page out of Google.
HTTP responses and validation
The server response affects how rules are interpreted. For Google:
- 2xx: the file is downloaded and processed.
- 3xx: the crawler follows several redirects before treating the file as unavailable.
- 4xx: apart from specific cases such as 429, the response is normally treated as though no crawl restrictions exist.
- 5xx or network errors: crawling may stop temporarily and a previously valid version may be used.
A technical check should cover:
- Availability:
/robots.txtresponds on the correct host, protocol and port. - Content: the server returns plain text rather than an HTML page or custom error.
- Syntax: groups, agents, paths and comments can be interpreted without ambiguity.
- Consistency: resources needed to render or understand pages intended for crawling are not blocked.
- Changes: updates may take time to appear because search engines temporarily cache the file.
The outcome should be checked for every relevant crawler because non-standard directives such as Crawl-delay are not handled in the same way by every search engine.
