{"id":20407,"date":"2020-01-29T16:47:46","date_gmt":"2020-01-29T16:47:46","guid":{"rendered":"https:\/\/www.arimetrics.com\/glosario-digital\/crawl-error"},"modified":"2026-09-24T20:40:38","modified_gmt":"2026-09-24T20:40:38","slug":"crawl-error","status":"publish","type":"encyclopedia","link":"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawl-error","title":{"rendered":"Crawl Error"},"content":{"rendered":"<p><img decoding=\"async\" class=\"boxpad alignright wp-image-23099 size-full\" src=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2021\/11\/crawl-error.jpg\" alt=\"Crawl Error\" width=\"300\" height=\"300\" srcset=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2021\/11\/crawl-error.jpg 300w, https:\/\/www.arimetrics.com\/wp-content\/uploads\/2021\/11\/crawl-error-150x150.jpg 150w\" sizes=\"(max-width: 300px) 100vw, 300px\" \/> <strong>Definition:<\/strong><\/p>\n<p>A <strong>crawl error<\/strong> is an issue that prevents or hinders a <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawler\">crawler<\/a> from accessing a URL or retrieving its resources as intended. It can affect one page, a group of URLs, or an entire site, depending on whether it originates in the address itself, the server, the network, or access rules.<\/p>\n<p>Crawling and <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/indexing\">indexing<\/a> are related but different processes. A URL can be crawled without being indexed, while an intentionally blocked page may present no error at all. Diagnosis needs to compare the observed response with the state the URL is expected to have.<\/p>\n\n<h2>How a crawl error occurs<\/h2>\n<p>The crawler requests a URL, receives a response, and attempts to follow the instructions and resources it needs. An issue exists when that path cannot be completed as intended. Common causes include:<\/p>\n<ul>\n<li><strong>DNS or connectivity:<\/strong> The domain does not resolve, the connection fails, or the response exceeds the available time.<\/li>\n<li><strong>Server errors:<\/strong> The URL returns a <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/error-500\">500 error<\/a> or another 5xx response that prevents content retrieval.<\/li>\n<li><strong>Missing URL:<\/strong> The address returns a <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/404-error\">404 error<\/a> or 410. This may be the correct response when the resource should no longer exist.<\/li>\n<li><strong>Access restrictions:<\/strong> The <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/robots-txt\">robots.txt<\/a> file, authentication, a firewall, or a 403 response prevents the bot from entering.<\/li>\n<li><strong>Faulty redirects:<\/strong> A <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/301-redirect\">redirect<\/a> creates a loop, adds too many hops, or leads to an inaccessible destination.<\/li>\n<li><strong>Soft 404:<\/strong> The page returns 200, but its content indicates that the resource does not exist or provides no useful answer.<\/li>\n<li><strong>Blocked required resources:<\/strong> Files needed to render or understand the page cannot be retrieved.<\/li>\n<\/ul>\n<h2>Expected errors and actual issues<\/h2>\n<p>Not every response other than 200 needs to be fixed. A permanently removed URL may correctly return 404 or 410, a private area may require authentication, and a deliberately blocked path may comply with the site&#8217;s access policy.<\/p>\n<p>An issue exists when the state conflicts with the intention: a page that should be available returns 5xx, an important URL is blocked, internal links point to missing addresses, or a redirect chain prevents access to the destination. A <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/noindex\">noindex<\/a> directive is not a crawl error either: it allows access but asks for the page to be excluded from the index.<\/p>\n<h2>How to identify the source<\/h2>\n<p><a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/google-search-console\">Google Search Console<\/a> distributes these signals across the Page indexing report, URL Inspection, and Crawl stats. Each report answers a different question: what state Google knows for a page, whether it can access a particular URL, and how crawling activity changes across the site.<\/p>\n<p>Diagnosis draws on several sources:<\/p>\n<ul>\n<li>the HTTP response and redirect chain;<\/li>\n<li>server logs showing actual bot requests;<\/li>\n<li>a site crawl to find broken links, loops, and isolated URLs;<\/li>\n<li>robots.txt, authentication, CDN, and firewall rules;<\/li>\n<li>the difference between an occasional response and a repeated pattern.<\/li>\n<\/ul>\n<h2>How to fix a crawl error<\/h2>\n<ol>\n<li><strong>Confirm the expected state:<\/strong> Determine whether the URL should exist, redirect, be blocked, or remain outside the index.<\/li>\n<li><strong>Reproduce the issue:<\/strong> Check the response, hops, and resources under the same conditions affecting the crawler.<\/li>\n<li><strong>Correct the cause:<\/strong> Address the server, DNS, link, redirect, access rule, or resource responsible for the failure.<\/li>\n<li><strong>Update references:<\/strong> Remove URLs that should no longer be crawled from internal links and discovery files.<\/li>\n<li><strong>Validate the result:<\/strong> Confirm that the final response matches the intention and that required content has not been blocked.<\/li>\n<li><strong>Observe recurrence:<\/strong> Review logs and reports for long enough to distinguish a stable recovery from an intermittent failure.<\/li>\n<\/ol>\n<h2>Impact on SEO and AI visibility<\/h2>\n<p>Persistent errors can hinder page discovery, delay updates, and consume crawling resources on paths that provide no value. Their effect depends on the scope, duration, and importance of the affected URLs; an isolated and expected 404 is not equivalent to a penalty.<\/p>\n<p>Accessibility is also part of <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/technical-seo\">technical SEO<\/a> and the technical foundation of GEO, but AI systems do not necessarily use the same bots or rules as a search engine. A site may allow search crawling while restricting other agents, or the reverse. Resolving an access issue improves the possibility of retrieving the content, but does not guarantee indexing, inclusion in a generative answer, or a citation.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A crawl error prevents a bot from accessing a URL. Learn its causes, how to diagnose and fix it, and its impact on SEO and AI visibility.<\/p>\n","protected":false},"author":6,"featured_media":0,"template":"","encyclopedia-tag":[1241],"class_list":["post-20407","encyclopedia","type-encyclopedia","status-publish","hentry","encyclopedia-tag-web-crawling"],"_links":{"self":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia\/20407","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia"}],"about":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/types\/encyclopedia"}],"author":[{"embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/users\/6"}],"wp:attachment":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/media?parent=20407"}],"wp:term":[{"taxonomy":"encyclopedia-tag","embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia-tag?post=20407"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}