{"id":20553,"date":"2020-01-30T13:00:32","date_gmt":"2020-01-30T13:00:32","guid":{"rendered":"https:\/\/www.arimetrics.com\/glosario-digital\/googlebot"},"modified":"2026-09-28T11:10:55","modified_gmt":"2026-09-28T11:10:55","slug":"googlebot","status":"publish","type":"encyclopedia","link":"https:\/\/www.arimetrics.com\/en\/digital-glossary\/googlebot","title":{"rendered":"GoogleBot"},"content":{"rendered":"<p><img decoding=\"async\" class=\"boxpad alignright wp-image-14415 size-full\" src=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/google_bot.jpg\" alt=\"Googlebot\" width=\"300\" height=\"300\" srcset=\"https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/google_bot.jpg 300w, https:\/\/www.arimetrics.com\/wp-content\/uploads\/2020\/01\/google_bot-150x150.jpg 150w\" sizes=\"(max-width: 300px) 100vw, 300px\" \/><strong>Definition:<\/strong><\/p>\n<p><strong>Googlebot<\/strong> is the generic name for the crawlers that Google Search uses to discover and retrieve web pages and other resources. Its two main variants are Googlebot Smartphone, which simulates a mobile device, and Googlebot Desktop, which simulates a desktop computer.<\/p>\n<p>Googlebot participates in crawling, but <strong>crawling a URL does not mean that it will be indexed or ranked<\/strong>. After the resource is retrieved, other Google systems process it, render it when necessary, and decide whether its content may enter the index.<\/p>\n\n<h2>How Googlebot works<\/h2>\n<p>Googlebot discovers URLs primarily through links found on pages that Google already knows. It can also obtain them from sitemaps and other sources used by Google. The general process includes several actions:<\/p>\n<ul>\n<li><strong>Discovery:<\/strong> Google learns about a new URL or detects that a known one may have changed.<\/li>\n<li><strong>Scheduling:<\/strong> Its systems decide which URL to request and when, based on crawl demand and the site&#8217;s capacity.<\/li>\n<li><strong>Request:<\/strong> Googlebot asks the server for the resource and receives its HTTP status code, headers, and accessible content.<\/li>\n<li><strong>Processing:<\/strong> Google analyzes the document and may request associated files, such as CSS, JavaScript, or images, to understand and render the page.<\/li>\n<li><strong>Revisit:<\/strong> The URL may be crawled again to check for changes, with no guaranteed fixed interval.<\/li>\n<\/ul>\n<p>Googlebot is a specific Google Search <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawler\">crawler<\/a>. It is not a single program that travels across the entire web in a linear sequence, but a distributed infrastructure that sends requests from many machines and IP addresses.<\/p>\n<h2>Googlebot Smartphone and Googlebot Desktop<\/h2>\n<p>Google distinguishes between two main variants for web crawling:<\/p>\n<ul>\n<li><strong>Googlebot Smartphone:<\/strong> It simulates someone using a mobile browser. This variant makes most requests because Google primarily indexes the mobile version of content.<\/li>\n<li><strong>Googlebot Desktop:<\/strong> It simulates a desktop browser and is used for a smaller proportion of requests.<\/li>\n<\/ul>\n<p>The variants identify themselves through different HTTP user-agent headers, but both use the Googlebot token in <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/robots-txt\">robots.txt<\/a>. Robots.txt therefore cannot allow one variant while selectively blocking the other.<\/p>\n<p>The former Deepbot and Freshbot distinction no longer represents Google&#8217;s current documentation or operation. These variants should not be confused with Googlebot-Image, Googlebot-Video, or the other crawlers and fetchers that Google uses for specific products and purposes.<\/p>\n<h2>Crawling, rendering, and indexing<\/h2>\n<p>Crawling is only one part of the Google Search process. The following stages should be distinguished:<\/p>\n<ul>\n<li><strong>Crawling:<\/strong> Googlebot requests a URL and downloads the part of the resource that it can process.<\/li>\n<li><strong>Rendering:<\/strong> Google&#8217;s systems may execute JavaScript and load required resources to observe generated content.<\/li>\n<li><strong>Indexing:<\/strong> Google analyzes the content, checks signals such as canonicalization, and decides whether to store a version in its index.<\/li>\n<li><strong>Ranking:<\/strong> For a query, Google orders the results it considers relevant through its ranking systems.<\/li>\n<\/ul>\n<p>A page may be crawled without entering the <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/indexing\">index<\/a>. Google may also know a URL even if its content has not been crawled recently. Making a page accessible to Googlebot removes technical barriers, but does not guarantee inclusion, update frequency, or rankings.<\/p>\n<h2>How to identify and verify Googlebot<\/h2>\n<p>Server logs make it possible to review requests attributed to Googlebot, including the URL, time, IP address, user-agent, and response code. However, <strong>a user-agent name can be spoofed<\/strong>, so finding the word Googlebot in a log is not sufficient.<\/p>\n<p>Google recommends verifying requests through one of these methods:<\/p>\n<ul>\n<li>Compare the source IP address with Google&#8217;s published IP ranges for its crawlers.<\/li>\n<li>Perform a reverse DNS lookup on the IP address and then use a forward lookup to confirm that the resulting hostname resolves to the same address.<\/li>\n<\/ul>\n<p>Verified hostnames for common crawlers end in googlebot.com or geo.googlebot.com. This check helps prevent mistakenly allowing, blocking, or rate limiting a bot that is only impersonating Googlebot.<\/p>\n<h2>How to control Googlebot access<\/h2>\n<p>Each mechanism addresses a different requirement:<\/p>\n<ul>\n<li><strong>Robots.txt:<\/strong> It manages which paths Googlebot may crawl, but does not protect confidential information or guarantee that a known URL disappears from search results.<\/li>\n<li><strong>Meta robots noindex or X-Robots-Tag:<\/strong> It requests that the resource not be included in Google Search. Googlebot must be able to crawl it to read the directive.<\/li>\n<li><strong>Authentication:<\/strong> A password or another access control prevents Googlebot and unauthorized users from retrieving private content.<\/li>\n<li><strong>HTTP status codes:<\/strong> Responses such as 404 or 410 indicate that a resource is no longer available; 429 and 5xx communicate different problems and may temporarily reduce crawling.<\/li>\n<\/ul>\n<p>Blocking a page in robots.txt while adding noindex can be counterproductive because Googlebot will be unable to access the page and detect noindex. The appropriate measure depends on whether the goal is to stop crawling, prevent indexing, remove a URL, or protect information.<\/p>\n<h2>How to analyze crawling problems<\/h2>\n<p>A diagnosis should combine several sources. Logs show the requests actually received by the server. <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/google-search-console\">Google Search Console<\/a> provides URL inspection, crawl statistics, and reports on indexing errors or exclusion reasons.<\/p>\n<p>Robots.txt, HTTP status codes, redirects, canonicals, sitemaps, internal links, and resources required for rendering should also be checked. Visit frequency does not depend merely on supposed authority. Google considers server response capacity, crawl demand, content changes, quality, and relevance, among other signals.<\/p>\n<p>Optimizing <a href=\"https:\/\/www.arimetrics.com\/en\/digital-glossary\/crawl-budget\">crawl budget<\/a> is particularly relevant for very large or rapidly changing sites and those with many discovered but unindexed URLs. On small or stable websites, maintaining an accurate sitemap and reviewing the indexing report is usually more useful than trying to force a particular Googlebot frequency.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Googlebot crawls pages for Google Search. Learn its variants, how crawling differs from indexing, how to verify it, and how to control access.<\/p>\n","protected":false},"author":6,"featured_media":80901,"template":"","encyclopedia-tag":[1241],"class_list":["post-20553","encyclopedia","type-encyclopedia","status-publish","has-post-thumbnail","hentry","encyclopedia-tag-web-crawling"],"_links":{"self":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia\/20553","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia"}],"about":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/types\/encyclopedia"}],"author":[{"embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/users\/6"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/media\/80901"}],"wp:attachment":[{"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/media?parent=20553"}],"wp:term":[{"taxonomy":"encyclopedia-tag","embeddable":true,"href":"https:\/\/www.arimetrics.com\/en\/wp-json\/wp\/v2\/encyclopedia-tag?post=20553"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}