{"id":25331,"date":"2024-12-09T00:00:00","date_gmt":"2024-12-09T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/crawling-december-http-caching\/"},"modified":"2024-12-09T00:00:00","modified_gmt":"2024-12-09T00:00:00","slug":"crawling-december-http-caching","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/crawling-december-http-caching\/","title":{"rendered":"Crawling December: HTTP caching"},"content":{"rendered":"<p class=\"gargardate\">Monday, December 9, 2024<\/p>\n<p>\nAllow us to cache, pretty please.\n<\/p>\n<p>\n  As the internet grew over the years, so did how much Google crawls. While Google&#8217;s crawling<br \/>\n  infrastructure supports heuristic caching mechanisms, in fact always had, the number of requests<br \/>\n  that can be returned from local caches has decreased: 10 years ago about 0.026% of the total<br \/>\n  fetches were cacheable, which is already not that impressive; today that number is 0.017%.\n<\/p>\n<h2 id=\"why-is-caching-important\" tabindex=\"-1\">Why is caching important?<\/h2>\n<p>\n  Caching is a critical piece of the large puzzle that is the internet. Caching allows pages to load<br \/>\n  lightning fast on revisits, it saves computing resources and thus also natural resources, and<br \/>\n  saves a tremendous amount of expensive bandwidth for both the clients and servers.\n<\/p>\n<p>\n  Especially if you have a large site with rarely-changing content under individual URLs, allowing<br \/>\n  caching locally may help your site be crawled more efficiently. Google&#8217;s crawling infrastructure<br \/>\n  supports heuristic HTTP caching as defined by the<br \/>\n  <a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9111.html\" class=\"external-link\">HTTP caching standard<\/a>,<br \/>\n  specifically through the <code>ETag<\/code> response- and <code>If-None-Match<\/code> request<br \/>\n  header, and the <code>Last-Modified<\/code> response- and <code>If-Modified-Since<\/code> request<br \/>\n  header.\n<\/p>\n<p>\n  We strongly recommend using <code>ETag<\/code> because it&#8217;s less prone to errors and mistakes (the<br \/>\n  value is not structured unlike the <code>Last-Modified<\/code> value). And, if you have the option,<br \/>\n  set them both: the internet will thank you. Maybe.\n<\/p>\n<p>\n  As for what you consider a change that requires clients to refresh their caches, that&#8217;s up to you.<br \/>\n  Our recommendation is that you require a cache refresh on significant changes to your content; if<br \/>\n  you only updated the copyright date at the bottom of your page, that&#8217;s probably not significant.\n<\/p>\n<h2 id=\"etag-and-if-none-match\" tabindex=\"-1\"><code>ETag<\/code> and <code>If-None-Match<\/code><\/h2>\n<p>\n  Google&#8217;s crawlers support <code>ETag<\/code> based conditional requests exactly as defined in the<br \/>\n  <a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9111.html\" class=\"external-link\">HTTP caching standard<\/a>.<br \/>\n  That is, to signal caching preference to Google&#8217;s crawlers, set the <code>Etag<\/code> value to any<br \/>\n  arbitrary ASCII string (usually a hash of the content or version number, but it could also be a<br \/>\n  piece of the \u03c0, up to you) unique to the representation of the content hosted by the accessed URL.<br \/>\n  For example, if you host different versions of the same content under the same URL (say, mobile<br \/>\n  and desktop version), each version could have its own unique <code>ETag<\/code> value.\n<\/p>\n<p>\n  Google&#8217;s crawlers that support caching will send the <code>ETag<\/code> value returned for a<br \/>\n  previous crawl of that URL in the <code>If-None-Match header<\/code>. If the <code>ETag<\/code><br \/>\n  value sent by the crawler matches the current value the server generated, your server should<br \/>\n  return an HTTP <code>304<\/code> (Not modified) status code with no HTTP body. This last bit, no<br \/>\n  HTTP body, is the important part for a couple reasons:\n<\/p>\n<ul>\n<li>\n    your server doesn&#8217;t have to spend compute resources on actually generating content; that is, you<br \/>\n    save money\n  <\/li>\n<li>\n    your server doesn&#8217;t have to transfer the HTTP body; that is, you save money\n  <\/li>\n<\/ul>\n<p>\n  On the client side, like a user&#8217;s browser or Googlebot, the content under that URL is retrieved<br \/>\n  from the client&#8217;s internal cache. Because there&#8217;s no data transfer involved, this happens<br \/>\n  lightning fast, making users happy and potentially saving some resources for them, too.\n<\/p>\n<h2 id=\"last-modified-and-if-modified-since\" tabindex=\"-1\"><code>Last-Modified<\/code> and <code>If-Modified-Since<\/code><\/h2>\n<p>\n  Similarly to <code>ETag<\/code>, Google&#8217;s crawlers support <code>Last-Modified based<\/code><br \/>\n  conditional requests, too, exactly as defined in the HTTP Caching standard. This works the same<br \/>\n  way as <code>ETag<\/code> from a semantic perspective &mdash; an identifier is used to decide<br \/>\n  whether the resource is cacheable &mdash;, and provides the same benefits as <code>ETag<\/code> on<br \/>\n  the clients&#8217; side.\n<\/p>\n<p>\n  We have but a couple recommendations if you&#8217;re using <code>Last-Modified<\/code> as a caching<br \/>\n  directive:\n<\/p>\n<ol>\n<li>\n    The date in the <code>Last-Modified<\/code> header must be formatted according to the<br \/>\n    <a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9110.html\" class=\"external-link\">HTTP standard<\/a>.<br \/>\n    To avoid parsing issues, we recommend using the following date format:<br \/>\n    &#8220;Weekday, <span>DD Mon YYYY HH:MM:SS<\/span> Timezone&#8221;. For example,<br \/>\n    &#8220;<span>Fri, 4 Sep 1998 19:15:56 GMT<\/span>&#8220;.\n  <\/li>\n<li>\n    While not required, consider also setting the <code>max-age<\/code> field of the<br \/>\n    <code>Cache-Control<\/code> header to help crawlers determine when to recrawl the specific URL.<br \/>\n    Set the value of the <code>max-age<\/code> field to the expected number of seconds the content<br \/>\n    will be unchanged. For example, <code>Cache-Control: max-age=94043<\/code>.\n  <\/li>\n<\/ol>\n<h2 id=\"examples\" tabindex=\"-1\">Examples<\/h2>\n<p>\n  If you&#8217;re like me, wrapping my head around how heuristic caching works is challenging, however<br \/>\n  showing an example of the chain of requests and responses seems to help me. Here are two chains<br \/>\n  &mdash; one for <code>ETag<\/code>\/<code>If-None-Match<\/code> and one for<br \/>\n  <code>Last-Modified<\/code>\/<code>If-Modified-Since<\/code> &mdash; to visualize how it&#8217;s supposed<br \/>\n  to work:\n<\/p>\n<table class=\"fixed vertical-rules full-width\">\n<tr>\n<th width=\"33%\"><\/th>\n<th><code>ETag<\/code>\/<code>If-None-Match<\/code><\/th>\n<th><code>Last-Modified<\/code>\/<code>If-Modified-Since<\/code><\/th>\n<\/tr>\n<tr>\n<td>\n      <strong>A server&#8217;s response to a crawl:<\/strong> This is the response from which a crawler can<br \/>\n      save the precondition header fields <code>ETag<\/code> and <code>Last-Modified<\/code>.\n    <\/td>\n<td>\n<div><\/div>\n<\/td>\n<td>\n<div><\/div>\n<\/td>\n<\/tr>\n<tr>\n<td>\n      <strong>Subsequent crawler conditional request:<\/strong> The conditional request is based on<br \/>\n      the precondition header values saved from a previous request. The values are sent back to the<br \/>\n      server for validation in the <code>If-None-Match<\/code> and <code>If-Modified-Since<\/code><br \/>\n      request headers.\n    <\/td>\n<td>\n<div><\/div>\n<\/td>\n<td>\n<div><\/div>\n<\/td>\n<\/tr>\n<tr>\n<td>\n      <strong>Server response to the conditional request:<\/strong> Since precondition header values<br \/>\n      sent by the crawler are validated on the server&#8217;s side, the server returns a <code>304<\/code><br \/>\n      HTTP status code (without an HTTP body) to the crawler. This will happen to every subsequent<br \/>\n      request until the preconditions fail to validate (the <code>ETag<\/code> or the<br \/>\n      <code>Last-Modified<\/code> date changes on the server&#8217;s side).\n    <\/td>\n<td>\n<div><\/div>\n<\/td>\n<td>\n<div><\/div>\n<\/td>\n<\/tr>\n<\/table>\n<p>\n  If you&#8217;re in the business of making your users happy and perhaps also want to potentially save a<br \/>\n  few bucks on your hosting bill, talk to your hosting or CMS provider, or your developers about how<br \/>\n  to enable HTTP caching for your site. If nothing else, your users will like you a bit more.\n<\/p>\n<p>\n  If you wanna chat about caching, head to your nearest<br \/>\n  <a href=\"https:\/\/goo.gle\/sc-forum\" class=\"external-link\">Search Central help community<\/a>, and if<br \/>\n   you have comments about how we&#8217;re caching, leave feedback on<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/google-crawlers\/#http-caching\">the documentation about caching<\/a><br \/>\n  that we published together with this blog post.\n<\/p>\n<p class=\"byline-author\">\n  Posted by <a href=\"https:\/\/developers.google.com\/search\/blog\/authors\/gary-illyes\">Gary Illyes<\/a>\n<\/p>\n<hr class=\"full-width\">\n<h2 id=\"want-to-learn-more-about-crawling-check-out-the-entire-crawling-december-series:\" tabindex=\"-1\">Want to learn more about crawling? Check out the entire Crawling December series:<\/h2>\n","protected":false},"excerpt":{"rendered":"<p>Monday, December 9, 2024 Allow us to cache, pretty please. As the internet grew over the years, so did how much Google crawls. While Google&#8217;s crawling infrastructure supports heuristic caching mechanisms, in fact always had, the number of requests that can be returned from local caches has decreased: 10 years ago about 0.026% of the [&hellip;]<\/p>\n","protected":false},"menu_order":79934,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-25331","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25331","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25331\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/27451"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=25331"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=25331"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=25331"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}