{"id":20279,"date":"2020-01-30T10:28:10","date_gmt":"2020-01-30T10:28:10","guid":{"rendered":"https:\/\/www.arimetrics.com\/glosario-digital\/crawler"},"modified":"2026-09-21T19:21:48","modified_gmt":"2026-09-21T19:21:48","slug":"crawler","status":"publish","type":"encyclopedia","link":"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawler","title":{"rendered":"Crawler"},"content":{"rendered":"<p><img decoding=\"async\" class=\"boxpad alignright size-full wp-image-14132\" style=\"margin-top:0;\" src=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/crawler.png\" alt=\"Crawler\" width=\"300\" height=\"300\" srcset=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/crawler.png 300w, https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/crawler-150x150.png 150w\" sizes=\"(max-width: 300px) 100vw, 300px\" \/><strong>Definition:<\/strong><\/p>\n<p>A <strong>crawler<\/strong> or web crawler is an automated program that discovers addresses, requests resources and follows links to collect information from pages and files. It is also known as a <em>spider<\/em>, <em>robot<\/em> or <em>bot<\/em> when it performs this journey systematically.<\/p>\n<p>Search engines use crawlers as the first stage of their systems, but crawling does not mean indexing or ranking. The program retrieves the resource and its technical signals; subsequent processes decide how to interpret it, whether it can enter the index and when it might appear as a result.<\/p>\n\n<h2>How a crawler works<\/h2>\n<p>The usual journey includes these stages:<\/p>\n<ol>\n<li><strong>URL discovery:<\/strong> obtain addresses from known links, a <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/sitemap\">sitemap<\/a>, a seed queue or another authorized source.<\/li>\n<li><strong>Scheduling:<\/strong> select which URL to request, when to request it and its priority according to objectives and resources.<\/li>\n<li><strong>Request:<\/strong> send a request to the server and receive an HTTP status code, headers and available content.<\/li>\n<li><strong>Processing:<\/strong> analyze the document, extract links, metadata and other elements and, when required, render resources or JavaScript.<\/li>\n<li><strong>Queue update:<\/strong> record the outcome, avoid unnecessary repetition and schedule new discoveries or revisits.<\/li>\n<\/ol>\n<p>A crawler does not necessarily visit every page. It may stop because of restrictions, errors, time or capacity limits, duplicates, low demand or selection decisions.<\/p>\n<h2>Policies that determine crawling<\/h2>\n<ul>\n<li><strong>Selection:<\/strong> determines which URLs enter the queue and which are rejected or postponed.<\/li>\n<li><strong>Revisit:<\/strong> decides when a resource should be requested again to check for changes.<\/li>\n<li><strong>Politeness:<\/strong> limits frequency and concurrency to avoid overloading the server.<\/li>\n<li><strong>Parallelization:<\/strong> distributes requests across processes, machines or domains while maintaining per-host controls.<\/li>\n<li><strong>Budget:<\/strong> allocates finite time and resources; in search engines, this availability relates to <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawl-budget\">crawl budget<\/a>.<\/li>\n<\/ul>\n<p>Behaviour depends on the crawler&#8217;s purpose. A search engine, web archive and audit tool may use different policies even though all of them request documents through HTTP.<\/p>\n<h2>Crawler types and uses<\/h2>\n<ul>\n<li><strong>Search engine crawlers:<\/strong> discover content for search systems; <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/googlebot\">Googlebot<\/a> is one specific example, not a synonym for every crawler.<\/li>\n<li><strong>Audit crawlers:<\/strong> visit a website to detect HTTP statuses, links, metadata, canonicals and technical <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/seo\">SEO<\/a> issues.<\/li>\n<li><strong>Archive crawlers:<\/strong> preserve copies and page changes to build historical collections.<\/li>\n<li><strong>Monitoring crawlers:<\/strong> check availability, modifications, prices or other defined properties.<\/li>\n<li><strong>Extraction crawlers:<\/strong> collect structured or semi-structured data for authorized research, comparison or analysis.<\/li>\n<\/ul>\n<p>The technical ability to request a page does not grant unlimited permission to copy or reuse its data. The scope must respect terms of service, privacy, intellectual property, server load and applicable law.<\/p>\n<h2>Crawling, indexing and ranking<\/h2>\n<ul>\n<li><strong>Discovery:<\/strong> the system knows a URL even if it has not requested it yet.<\/li>\n<li><strong>Crawling:<\/strong> the crawler obtains the resource and observes its accessible response and content.<\/li>\n<li><strong><a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/indexing\">Indexing<\/a>:<\/strong> the search engine processes the page, evaluates its content and decides whether to store a version in its index.<\/li>\n<li><strong>Canonicalization:<\/strong> groups similar pages and selects a representative version using signals such as <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/rel-canonical\">rel canonical<\/a>.<\/li>\n<li><strong>Ranking:<\/strong> orders results for a particular query after considering relevance, quality and context.<\/li>\n<\/ul>\n<p>A crawled URL may not be indexed, and a known URL may appear in a limited form without its content having been crawled. Making a page accessible to a crawler does not guarantee rankings either.<\/p>\n<h2>How crawler access is controlled<\/h2>\n<ul>\n<li><strong><a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/robots-txt\">Robots.txt<\/a>:<\/strong> communicates crawling rules to compatible bots, but it does not protect private information or guarantee that every bot will comply.<\/li>\n<li><strong>Authentication:<\/strong> a password or access control prevents unauthorized visitors from retrieving the resource.<\/li>\n<li><strong><code>noindex<\/code>:<\/strong> asks compatible search engines not to add the page to their index, but they normally need to crawl it to read the directive.<\/li>\n<li><strong>HTTP statuses:<\/strong> codes such as 401, 403, 404, 410, 429 and 5xx communicate different situations and affect revisits.<\/li>\n<li><strong>Server limits:<\/strong> rate limiting, firewalls and other measures can control requests, although an incorrect configuration may also block legitimate bots.<\/li>\n<\/ul>\n<p>Blocking a URL in <code>robots.txt<\/code> manages crawling, not confidentiality or necessarily the presence of the address as a known URL in an index.<\/p>\n<h2>Crawling analysis and issues<\/h2>\n<ul>\n<li><strong>Server logs:<\/strong> show which user agent requested each URL, when it did so and which response it received.<\/li>\n<li><strong>Bot verification:<\/strong> a user agent name can be spoofed; search engines document validation methods using IP and DNS.<\/li>\n<li><strong>Search reports:<\/strong> <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/google-search-console\">Google Search Console<\/a> provides crawling and indexing statuses for a verified property.<\/li>\n<li><strong>Controlled audits:<\/strong> a first-party crawler can reproduce journeys, check links and compare responses without assuming that it exactly imitates a search engine.<\/li>\n<li><strong>Combined diagnosis:<\/strong> architecture, internal links, sitemaps, robots rules, HTTP statuses, canonicals, performance and rendering should be analyzed as one system.<\/li>\n<\/ul>\n<p>Common problems include loops, infinite calendars, combinable parameters, redirect chains, broken links, slow servers and content that depends on blocked resources. Correcting them improves crawling efficiency but does not replace subsequent content evaluation.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>What a crawler or web spider is, how it discovers and processes URLs, the main crawler types, and the differences between crawling, indexing and ranking.<\/p>\n","protected":false},"author":6,"featured_media":0,"template":"","encyclopedia-tag":[1241],"class_list":["post-20279","encyclopedia","type-encyclopedia","status-publish","hentry","encyclopedia-tag-web-crawling"],"_links":{"self":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia\/20279","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia"}],"about":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/types\/encyclopedia"}],"author":[{"embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/users\/6"}],"wp:attachment":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/media?parent=20279"}],"wp:term":[{"taxonomy":"encyclopedia-tag","embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia-tag?post=20279"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}