{"id":23846,"date":"2006-11-10T00:00:00","date_gmt":"2006-11-10T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/the-number-of-pages-googlebot-crawls\/"},"modified":"2006-11-10T00:00:00","modified_gmt":"2006-11-10T00:00:00","slug":"the-number-of-pages-googlebot-crawls","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/the-number-of-pages-googlebot-crawls\/","title":{"rendered":"The number of pages Googlebot crawls"},"content":{"rendered":"<p class=\"gargardate\">Friday, November 10, 2006<\/p>\n<p>\n  The <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/googlebot-activity-reports\/\">Googlebot activity reports<\/a><br \/>\n  in Webmaster Tools show you the number of pages of your site Googlebot has crawled over the last<br \/>\n  90 days. We&#8217;ve seen some of you asking why this number might be higher than the total number of<br \/>\n  pages on your sites.\n<\/p>\n<p><img decoding=\"async\" alt=\"number of pages crawled statistics in webmaster tools\" src=\"https:\/\/alienroad.com\/wp-content\/uploads\/kb-gorsel\/g-1a8df45811cc.png\" loading=\"lazy\"><\/p>\n<p>Googlebot crawls pages of your site based on a number of things including:<\/p>\n<ul>\n<li>pages it already knows about<\/li>\n<li>links from other web pages (within your site and on other sites)<\/li>\n<li>pages listed in your Sitemap file<\/li>\n<\/ul>\n<p>\n  More specifically, Googlebot doesn&#8217;t access pages, it accesses URLs. And the same page can often<br \/>\n  be accessed via several URLs. Consider the home page of a site that can be accessed from the<br \/>\n  following four URLs:\n<\/p>\n<ul>\n<li>https:\/\/www.example.com\/<\/li>\n<li>https:\/\/www.example.com\/index.html<\/li>\n<li>https:\/\/example.com<\/li>\n<li>https:\/\/example.com\/index.html<\/li>\n<\/ul>\n<p>\n  Although all URLs lead to the same page, all four URLs may be used in links to the page. When<br \/>\n  Googlebot follows these links, a count of four is added to the activity report.\n<\/p>\n<p>\n  Many other scenarios can lead to multiple URLs for the same page. For instance, a page may have<br \/>\n  several named anchors, such as:\n<\/p>\n<ul>\n<li>https:\/\/www.example.com\/mypage.html#heading1<\/li>\n<li>https:\/\/www.example.com\/mypage.html#heading2<\/li>\n<li>https:\/\/www.example.com\/mypage.html#heading3<\/li>\n<\/ul>\n<p>And dynamically generated pages often can be reached by multiple URLs, such as:<\/p>\n<ul>\n<li>https:\/\/www.example.com\/furniture?type=chair&amp;brand=123<\/li>\n<li>https:\/\/www.example.com\/hotbuys?type=chair&amp;brand=123<\/li>\n<\/ul>\n<p>\n  As you can see, when you consider that each page on your site might have multiple URLs that lead<br \/>\n  to it, the number of URLs that Googlebot crawls can be considerably higher than the number of<br \/>\n  total pages for your site.\n<\/p>\n<p>\n  Of course, you (and we) only want one version of the URL to be returned in the search results.<br \/>\n  Not to worry&mdash;this is exactly what happens. Our algorithms selects a version to include, and<br \/>\n  you can provide input on this selection process.\n<\/p>\n<h2 id=\"redirect-to-the-preferred-version-of-the-url\" tabindex=\"-1\">Redirect to the preferred version of the URL<\/h2>\n<p>\n  You can do this using<br \/>\n  <a href=\"https:\/\/en.wikipedia.org\/wiki\/URL_redirection\" class=\"external-link\"><code>301 (permanent)<\/code> redirect<\/a>.<br \/>\n  In the first example that shows four URLs that point to a site&#8217;s home page, you may want to<br \/>\n  redirect index.html to www.example.com\/. And you may want to redirect example.com to<br \/>\n  www.example.com so that any URLs that begin with one version are redirected to the other version.<br \/>\n  Note that you can do this latter redirect with the<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/setting-the-preferred-domain\/\">Preferred Domain feature<\/a> in webmaster<br \/>\n  tools. (If you also use a <code>301<\/code> redirect, make sure that this redirect matches what you<br \/>\n  set for the preferred domain.)\n<\/p>\n<h2 id=\"block-the-non-preferred-versions-of-a-url-with-a-robots.txt-file\" tabindex=\"-1\">Block the non-preferred versions of a URL with a robots.txt file<\/h2>\n<p>\n  For dynamically generated pages, you may want to block the non-preferred version using<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">pattern matching<\/a><br \/>\n  in your robots.txt file. (Note that not all search engines support pattern matching, so check<br \/>\n  the guidelines for each search engine bot you&#8217;re interested in.) For instance, in the third<br \/>\n  example that shows two URLs that point to a page about the chairs available from brand 123, the<br \/>\n  &#8220;hotbuys&#8221; section rotates periodically and the content is always available from a primary and<br \/>\n  permanent location. If that case, you may want to index the first version, and block the<br \/>\n  &#8220;<span>hotbuys<\/span>&#8221; version. To do this, add the following to your robots.txt<br \/>\n  file:\n<\/p>\n<div><\/div>\n<p>\n  To ensure that this rule will actually block and allow what you intend, use the robots.txt<br \/>\n  analysis tool in Webmaster Tools. Just add this rule to the robots.txt section on that<br \/>\n  page, list the URLs you want to check in the &#8220;Test URLs&#8221; section and click the Check button. For<br \/>\n  this example, you&#8217;d see a result like this:\n<\/p>\n<p><img decoding=\"async\" alt=\"robots.txt tester feature in webmaster tools\" src=\"https:\/\/alienroad.com\/wp-content\/uploads\/kb-gorsel\/g-a8fbff694725.png\" loading=\"lazy\"><\/p>\n<p>\n  Don&#8217;t worry about links to anchors, because while Googlebot will crawl each link, our algorithms<br \/>\n  will index the URL without the anchor.\n<\/p>\n<p>\n  And if you don&#8217;t provide input such as that described above, our algorithms do a really good job<br \/>\n  of picking a version to show in the search results.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Friday, November 10, 2006 The Googlebot activity reports in Webmaster Tools show you the number of pages of your site Googlebot has crawled over the last 90 days. We&#8217;ve seen some of you asking why this number might be higher than the total number of pages on your sites. Googlebot crawls pages of your site [&hellip;]<\/p>\n","protected":false},"menu_order":86538,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23846","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23846","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23846\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26478"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23846"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23846"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23846"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}