{"id":21463,"date":"2020-01-29T15:15:20","date_gmt":"2020-01-29T15:15:20","guid":{"rendered":"https:\/\/www.arimetrics.com\/glosario-digital\/robots-txt"},"modified":"2026-09-22T08:41:55","modified_gmt":"2026-09-22T08:41:55","slug":"robots-txt","status":"publish","type":"encyclopedia","link":"https:\/\/www.arimetrics.com\/en\/digital-glossary\/robots-txt","title":{"rendered":"Robots.txt"},"content":{"rendered":"<p><img decoding=\"async\" class=\"boxpad alignright size-full wp-image-14858\" style=\"margin-top:0;\" src=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/robots_txt.png\" alt=\"Robots.txt\" width=\"300\" height=\"300\" srcset=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/robots_txt.png 300w, https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/robots_txt-150x150.png 150w\" sizes=\"(max-width: 300px) 100vw, 300px\" \/><strong>Definition:<\/strong><\/p>\n<p><strong>Robots.txt<\/strong> is a text file located at a site&#8217;s root that tells crawlers which paths they may request and which they should not. Its rules form part of the Robots Exclusion Protocol and apply to <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawler\">crawlers<\/a> that recognise them.<\/p>\n<p>The file controls automated access for crawling. It is not linked from the HTML, does not protect private content and does not guarantee that a URL will disappear from search results.<\/p>\n\n<h2>How robots.txt works<\/h2>\n<p>A crawler requests the file before fetching other URLs from the site:<\/p>\n<ol>\n<li><strong>Request:<\/strong> the bot attempts to download <code>\/robots.txt<\/code> from the host it intends to crawl.<\/li>\n<li><strong>Identification:<\/strong> it compares its identifier with the groups defined through <code>User-agent<\/code>.<\/li>\n<li><strong>Rule selection:<\/strong> it applies the <code>Allow<\/code> and <code>Disallow<\/code> directives in the most specific group that matches it.<\/li>\n<li><strong>Path matching:<\/strong> it compares each URL with the declared paths and uses the most specific match.<\/li>\n<li><strong>Crawl decision:<\/strong> it requests or avoids the URL according to the result and the behaviour implemented by that crawler.<\/li>\n<\/ol>\n<p>The file is public and its instructions are not an authorisation mechanism. A malicious bot can ignore them, and anyone can inspect the paths listed in it.<\/p>\n<h2>Location and scope<\/h2>\n<p>The <a href=\"https:\/\/www.rfc-editor.org\/info\/rfc9309\/\" target=\"_blank\" rel=\"noopener\">REP standard<\/a> specifies that the file must be named <code>robots.txt<\/code>, use lowercase letters and be available at the top-level path of the service.<\/p>\n<ul>\n<li><strong>Site root:<\/strong> for <code>https:\/\/www.example.com\/<\/code>, the location is <code>https:\/\/www.example.com\/<wbr>robots.txt<\/code>.<\/li>\n<li><strong>Host:<\/strong> rules on <code>www.example.com<\/code> do not automatically apply to <code>shop.example.com<\/code>.<\/li>\n<li><strong>Protocol:<\/strong> a file served over HTTPS does not by itself define the rules for the HTTP version.<\/li>\n<li><strong>Port:<\/strong> a non-standard port has its own scope.<\/li>\n<li><strong>Format:<\/strong> it must be served as UTF-8 plain text; a file placed inside a folder does not control the whole site.<\/li>\n<\/ul>\n<p>The path is case-sensitive. <code>\/robots.txt<\/code> is the expected location, whereas <code>\/Robots.txt<\/code> may be treated as a different resource.<\/p>\n<h2>Directives and syntax<\/h2>\n<p>Lines are arranged into groups containing one or more agents and their rules:<\/p>\n<ul>\n<li><strong><code>User-agent<\/code>:<\/strong> identifies the crawler addressed by the group. The value <code>*<\/code> represents agents without a more specific group.<\/li>\n<li><strong><code>Disallow<\/code>:<\/strong> identifies a path that the agent should not request. An empty value does not block any path.<\/li>\n<li><strong><code>Allow<\/code>:<\/strong> permits a specific path within a broader blocked path.<\/li>\n<li><strong><code>Sitemap<\/code>:<\/strong> declares the absolute URL of a <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/sitemap\">sitemap<\/a>. This line does not belong to an agent group.<\/li>\n<li><strong>Comments:<\/strong> the <code>#<\/code> character causes the rest of the line to be interpreted as a comment.<\/li>\n<\/ul>\n<p><code>Allow<\/code> and <code>Disallow<\/code> paths start with <code>\/<\/code>. Specificity determines which rule applies when several rules match. Google&#8217;s <a href=\"https:\/\/developers.google.com\/crawling\/docs\/robots-txt\/robots-txt-spec\" target=\"_blank\" rel=\"noopener\">robots.txt syntax documentation<\/a> also supports <code>*<\/code> for a sequence of characters and <code>$<\/code> to mark the end of a URL.<\/p>\n<h2>Robots.txt examples<\/h2>\n<h3>Allow all crawling<\/h3>\n<pre><code class=\"language-text\">User-agent: *\nDisallow:\n<\/code><\/pre>\n<p>Leaving the path after <code>Disallow<\/code> empty imposes no restriction.<\/p>\n<h3>Block a directory<\/h3>\n<pre><code class=\"language-text\">User-agent: *\nDisallow: \/private\/\n<\/code><\/pre>\n<p>The rule asks compatible agents not to crawl URLs whose paths begin with <code>\/private\/<\/code>.<\/p>\n<h3>Allow an exception inside a blocked directory<\/h3>\n<pre><code class=\"language-text\">User-agent: *\nDisallow: \/resources\/\nAllow: \/resources\/public\/\n<\/code><\/pre>\n<p>The more specific match allows <code>\/resources\/public\/<\/code> to be crawled even though its parent directory is blocked.<\/p>\n<h3>Block the entire site and declare a sitemap<\/h3>\n<pre><code class=\"language-text\">User-agent: *\nDisallow: \/\n\nSitemap: https:\/\/www.example.com\/sitemap.xml\n<\/code><\/pre>\n<p><code>Disallow: \/<\/code> requests that all crawling be blocked for that group. The <code>Sitemap<\/code> line can still be present, but it does not override the block.<\/p>\n<h2>Crawling, indexing and security<\/h2>\n<p>Robots.txt should be distinguished from other controls:<\/p>\n<ul>\n<li><strong>Crawling:<\/strong> the rules govern whether a compatible agent may request the content at a path.<\/li>\n<li><strong><a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/indexing\">Indexing<\/a>:<\/strong> a blocked URL may still appear in results if a search engine discovers it through links or other sources.<\/li>\n<li><strong><a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/noindex\">Noindex<\/a>:<\/strong> this directive is placed in a meta tag or HTTP header to request exclusion from the index. The search engine needs to crawl the URL to see it.<\/li>\n<li><strong>Privacy:<\/strong> a path listed in robots.txt remains public and directly accessible. Confidential content requires authentication or access controls.<\/li>\n<li><strong>Removal:<\/strong> blocking crawling does not erase a known URL or remove copies stored by other systems.<\/li>\n<\/ul>\n<p>Google&#8217;s <a href=\"https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/robots\/intro\" target=\"_blank\" rel=\"noopener\">robots.txt guide<\/a> explicitly states that the file manages crawl traffic and is not a mechanism for keeping a page out of Google.<\/p>\n<h2>HTTP responses and validation<\/h2>\n<p>The server response affects how rules are interpreted. For Google:<\/p>\n<ul>\n<li><strong>2xx:<\/strong> the file is downloaded and processed.<\/li>\n<li><strong>3xx:<\/strong> the crawler follows several redirects before treating the file as unavailable.<\/li>\n<li><strong>4xx:<\/strong> apart from specific cases such as 429, the response is normally treated as though no crawl restrictions exist.<\/li>\n<li><strong>5xx or network errors:<\/strong> crawling may stop temporarily and a previously valid version may be used.<\/li>\n<\/ul>\n<p>A technical check should cover:<\/p>\n<ul>\n<li><strong>Availability:<\/strong> <code>\/robots.txt<\/code> responds on the correct host, protocol and port.<\/li>\n<li><strong>Content:<\/strong> the server returns plain text rather than an HTML page or custom error.<\/li>\n<li><strong>Syntax:<\/strong> groups, agents, paths and comments can be interpreted without ambiguity.<\/li>\n<li><strong>Consistency:<\/strong> resources needed to render or understand pages intended for crawling are not blocked.<\/li>\n<li><strong>Changes:<\/strong> updates may take time to appear because search engines temporarily cache the file.<\/li>\n<\/ul>\n<p>The outcome should be checked for every relevant crawler because non-standard directives such as <code>Crawl-delay<\/code> are not handled in the same way by every search engine.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>What a robots.txt file is, where it belongs, how its directives control crawler access, and why blocking crawling does not always prevent indexing.<\/p>\n","protected":false},"author":7,"featured_media":0,"template":"","encyclopedia-tag":[1241],"class_list":["post-21463","encyclopedia","type-encyclopedia","status-publish","hentry","encyclopedia-tag-web-crawling"],"_links":{"self":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia\/21463","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia"}],"about":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/types\/encyclopedia"}],"author":[{"embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/users\/7"}],"wp:attachment":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/media?parent=21463"}],"wp:term":[{"taxonomy":"encyclopedia-tag","embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia-tag?post=21463"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}