{"id":23958,"date":"2008-03-06T00:00:00","date_gmt":"2008-03-06T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/first-date-with-the-googlebot-headers-and-compression\/"},"modified":"2008-03-06T00:00:00","modified_gmt":"2008-03-06T00:00:00","slug":"first-date-with-the-googlebot-headers-and-compression","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/first-date-with-the-googlebot-headers-and-compression\/","title":{"rendered":"First date with the Googlebot: Headers and compression"},"content":{"rendered":"<aside class=\"key-point\">It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore).<\/aside>\n<p class=\"gargardate\">Thursday, March 06, 2008<\/p>\n<p><img decoding=\"async\" class=\"attempt-right\" alt=\"googlebot with flowers\" src=\"https:\/\/alienroad.com\/wp-content\/uploads\/kb-gorsel\/g-92c4e4b0412c.jpg\" loading=\"lazy\"><\/p>\n<p>\n  <b>Name\/User-Agent<\/b>: Googlebot<br \/>\n  <b>IP Address<\/b>:<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/verify-requests-from-google-crawlers-and-fetchers\/\">Learn how to verify Googlebot<\/a> <br \/>\n  <b>Looking For<\/b>: Websites with unique and compelling content<br \/>\n  <b>Major Turn Off<\/b>: Violations of the<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/search-essentials-overview\/\">Webmaster Guidelines<\/a><br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/googlebot\/\">Googlebot<\/a> &mdash;what a dreamboat. It&#8217;s<br \/>\n  like they know us <code>&lt;head&gt;<\/code>, <code>&lt;body&gt;<\/code>, and soul. They&#8217;re probably<br \/>\n  not looking for anything exclusive; they see billions of other sites (though we share our data<br \/>\n  with other bots as well), but tonight we&#8217;ll really get to know each other as website and crawler.\n<\/p>\n<p>\n  I know, it&#8217;s never good to over-analyze a first date. We&#8217;re going to get to know Googlebot a bit<br \/>\n  more slowly, in a series of posts:\n<\/p>\n<ol>\n<li>\n    Our first date (tonight!): Headers Googlebot sends, file formats they &#8220;notice,&#8221; whether it&#8217;s<br \/>\n    better to compress data\n  <\/li>\n<li>\n    Judging their response: Response codes (<code>301<\/code>, <code>302<\/code>), how they handle<br \/>\n    redirects and <code>If-Modified-Since<\/code>\n  <\/li>\n<li>\n    Next steps: Following links, having them crawl faster or slower (so they don&#8217;t come on too<br \/>\n    strong)\n  <\/li>\n<\/ol>\n<p>And tonight is just the first date&#8230;<\/p>\n<hr>\n<p><b>Googlebot:<\/b> <span>ACK<\/span><\/p>\n<p><b>Website:<\/b> Googlebot, you&#8217;re here!<\/p>\n<p><b>Googlebot:<\/b> I am.<\/p>\n<div><\/div>\n<p>\n  <b>Website:<\/b> Those headers are so flashy!  Would you crawl with the same headers if my site<br \/>\n  were in the U.S., Asia or Europe? Do you ever use different headers?\n<\/p>\n<p>\n  <b>Googlebot:<\/b> My headers are typically consistent world-wide. I&#8217;m trying to see what a page<br \/>\n  looks like for the default language and settings for the site. Sometimes the<br \/>\n  <code>User-Agent<\/code> is different, for instance AdSense fetches use<br \/>\n  <code>Mediapartners-Google<\/code>:\n<\/p>\n<div><\/div>\n<p>Or for image search:<\/p>\n<div><\/div>\n<p>\n  Wireless fetches often have carrier-specific user agents, whereas Google Reader RSS fetches<br \/>\n  include extra info such as number of subscribers.\n<\/p>\n<p>\n  I usually avoid cookies (so no <code>Cookie:<\/code> header) since I don&#8217;t want the content<br \/>\n  affected too much by session-specific info. And, if a server uses a session id in a dynamic URL<br \/>\n  rather than a cookie, I can usually figure this out, so that I don&#8217;t end up crawling your same<br \/>\n  page a million times with a million different session ids.\n<\/p>\n<p>\n  <b>Website:<\/b> I&#8217;m very complex. I have many file types. Your headers say<br \/>\n  <code>Accept: *\/*<\/code>. Do you index all URLs or are certain file extensions automatically<br \/>\n  filtered?\n<\/p>\n<p><b>Googlebot:<\/b> That depends on what I&#8217;m looking for. If I&#8217;m indexing for regular web search,<br \/>\n  and I see links to MP3s and videos, I probably won&#8217;t download those. Similarly, if I see a JPG, I<br \/>\n  will treat it differently than an HTML or PDF link. For instance, JPG is much less likely to<br \/>\n  change frequently than HTML, so I will check the JPG for changes less often to save bandwidth.<br \/>\n  Meanwhile, if I&#8217;m looking for links as Google Scholar, I&#8217;m going to be far more interested in the<br \/>\n  PDF article than the JPG file. Downloading doodles (like JPGs) and videos of skateboarding dogs<br \/>\n  is distracting for a scholar&mdash;do you agree?\n<\/p>\n<p>\n  <b>Website:<\/b> Yes, they can be distracting. I&#8217;m in awe of your dedication. I love doodles (JPGs)<br \/>\n  and find them hard to resist.\n<\/p>\n<p>\n  <b>Googlebot:<\/b> Me, too; I&#8217;m not always so scholarly. When I crawl for image search, I&#8217;m very<br \/>\n  interested in JPGs. And for news, I&#8217;m mostly looking at HTML and nearby images.\n<\/p>\n<p>\n  There are also plenty of extensions (exe, dll, zip, dmg&#8230;), that tend to be big and less useful<br \/>\n  for a search engine.\n<\/p>\n<p>\n  <b>Website:<\/b> If you saw my URL, <code>https:\/\/www.example.com\/page1.LOL111<\/code>, would you<br \/>\n  (whimper whimper) reject it just because it contains an unknown file extension?\n<\/p>\n<p>\n  <b>Googlebot:<\/b> Website, let me give a bit more background. After actually downloading a file, I<br \/>\n  use the <code>Content-Type<\/code> header to check whether it really is HTML, an image, text, or<br \/>\n  something else. If it&#8217;s a special data type like a PDF file, Word document, or Excel spreadsheet, I<br \/>\n  &#8216;ll make sure it&#8217;s in the valid format and extract the text content. Maybe it has a virus; you<br \/>\n  never know. If the document or data type is really garbled, there&#8217;s usually not much to do besides<br \/>\n  discard the content.\n<\/p>\n<p>\n  So, if I&#8217;m crawling <code>https:\/\/www.example.com\/page1.LOL111<\/code> with an unknown file<br \/>\n  extension, it&#8217;s likely that I would start to download it. If I can&#8217;t figure out the content type<br \/>\n  from the header, or it&#8217;s a format that we don&#8217;t index (for example, mp3), then it&#8217;ll be put aside.<br \/>\n  Otherwise, we proceed indexing the file.\n<\/p>\n<p>\n  <b>Website:<\/b> My apologies for scrutinizing your style, Googlebot, but I noticed your<br \/>\n  <code>Accept-Encoding<\/code> headers say:\n<\/p>\n<div><\/div>\n<p>Can you explain these headers to me?<\/p>\n<p>\n  <b>Googlebot:<\/b> Sure. All major search engines and web browsers support gzip compression for<br \/>\n  content to save bandwidth. Other entries that you might see here include <code>x-gzip<\/code> (the<br \/>\n  same as <code>gzip<\/code>), <code>deflate<\/code> (which we also support), and<br \/>\n  <code>identity<\/code> (none).\n<\/p>\n<p>\n  <b>Website:<\/b> Can you talk more about file compression and<br \/>\n  <code>Accept-Encoding: gzip,deflate<\/code>? Many of my URLs consist of big Flash files and<br \/>\n  stunning images, not just HTML. Would it help you to crawl faster if I compressed my larger files?\n<\/p>\n<p>\n  <b>Googlebot:<\/b> There&#8217;s not a simple answer to this question. First of all, many file formats,<br \/>\n  such as swf (Flash), jpg, png, gif, and pdf are already compressed (there are also specialized<br \/>\n  Flash optimizers).\n<\/p>\n<p>\n  <b>Website:<\/b> Perhaps I&#8217;ve been compressing my Flash files and I didn&#8217;t even know? I&#8217;m obviously<br \/>\n  very efficient.\n<\/p>\n<p>\n  <b>Googlebot:<\/b> Both Apache and IIS have options to enable gzip and deflate compression, though<br \/>\n  there&#8217;s a CPU cost involved for the bandwidth saved. Typically, it&#8217;s only enabled for easily<br \/>\n  compressible text HTML\/CSS\/PHP content. And it only gets used if the user&#8217;s browser or I (a search<br \/>\n  engine crawler) allow it. Personally, I prefer <code>gzip<\/code> over <code>deflate<\/code>. Gzip<br \/>\n  is a slightly more robust encoding&mdash;there is consistently a checksum and a full header,<br \/>\n  giving me less guess-work than with deflate. Otherwise they&#8217;re very similar compression<br \/>\n  algorithms.\n<\/p>\n<p>\n  If you have some spare CPU on your servers, it might be worth experimenting with compression<br \/>\n  (links:<br \/>\n  <a href=\"https:\/\/www.sitepoint.com\/article\/web-output-mod_gzip-apache\" class=\"external-link\">Apache<\/a>,<br \/>\n  <a href=\"https:\/\/www.microsoft.com\/technet\/prodtechnol\/WindowsServer2003\/Library\/IIS\/502ef631-3695-4616-b268-cbe7cf1351ce.mspx?mfr=true\" class=\"external-link\">IIS<\/a>).<br \/>\n  But, if you&#8217;re serving dynamic content and your servers are already heavily CPU loaded, you might<br \/>\n  want to hold off.\n<\/p>\n<p>\n  <b>Website:<\/b> Great information. I&#8217;m really glad you came tonight&mdash;thank goodness my<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">robots.txt<\/a> allowed it. That file can be like an<br \/>\n  over-protective parent!\n<\/p>\n<p>\n  <b>Googlebot:<\/b> Ah yes; meeting the parents, the robots.txt. I&#8217;ve met plenty of intense ones.<br \/>\n  Some are really just HTML error pages rather than valid robots.txt. Some have infinite redirects<br \/>\n  all over the place, maybe to totally unrelated sites, while others are just huge and have<br \/>\n  thousands of different URLs listed individually. Here&#8217;s one unfortunate pattern. The site is<br \/>\n  normally eager for me to crawl:\n<\/p>\n<div><\/div>\n<p>\n  Then, during a peak time with high user traffic, the site switches the robots.txt to something<br \/>\n  restrictive:\n<\/p>\n<div><\/div>\n<p>\n  The problem with the above robots.txt file-swapping is that once I see the restrictive robots.txt,<br \/>\n  I may have to start throwing away content I&#8217;ve already crawled in the index. And then I have to<br \/>\n  recrawl a lot of content once I&#8217;m allowed to crawl the site again. At least a 503 response code<br \/>\n  would&#8217;ve been temporary.\n<\/p>\n<p>\n  I typically only re-check robots.txt once a day (otherwise on many virtual hosting sites, I&#8217;d be<br \/>\n  spending a large fraction of my fetches just getting robots.txt, and no date wants to &#8220;meet the<br \/>\n  parents&#8221; that often). For webmasters, trying to control crawl rate through robots.txt swapping<br \/>\n  usually backfires. It&#8217;s better to<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/answer\/48620\" class=\"external-link\">set the rate to &#8220;slower&#8221;<\/a><br \/>\n  in Webmaster Tools.\n<\/p>\n<p>\n  <b>Googlebot:<\/b> Website, thanks for all of your questions, you&#8217;ve been wonderful, but I&#8217;m going<br \/>\n  to have to say &#8220;FIN, my love.&#8221;\n<\/p>\n<p>\n  <b>Website:<\/b> Oh, Googlebot&#8230; <span>ACK\/FIN<\/span>.<br \/>\n  <span>\ud83d\ude42<\/span>\n<\/p>\n<hr>\n<p class=\"byline-author\">Written by <a href=\"https:\/\/developers.google.com\/search\/blog\/authors\/maile-ohye\">Maile Ohye<\/a> as the website, Jeremy Lilley as the Googlebot<\/p>\n","protected":false},"excerpt":{"rendered":"<p>It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore). Thursday, March 06, 2008 Name\/User-Agent: Googlebot IP Address: Learn how to verify Googlebot Looking For: Websites with unique and compelling content Major Turn Off: [&hellip;]<\/p>\n","protected":false},"menu_order":86056,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23958","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23958","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23958\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26577"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23958"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23958"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23958"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}