{"id":25343,"date":"2025-03-28T00:00:00","date_gmt":"2025-03-28T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-refresher-future-proof-robots-exclusion-protocol\/"},"modified":"2025-03-28T00:00:00","modified_gmt":"2025-03-28T00:00:00","slug":"robots-refresher-future-proof-robots-exclusion-protocol","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-refresher-future-proof-robots-exclusion-protocol\/","title":{"rendered":"Robots Refresher: Future-proof Robots Exclusion Protocol"},"content":{"rendered":"<p class=\"gargardate\">Friday, March 28, 2025<\/p>\n<p>\n  In the previous posts about the Robots Exclusion Protocol (REP), we explored what you can already<br \/>\n  do with its various components &mdash; namely robots.txt and the URI level controls.<br \/>\n  In this post we will explore how the REP can play a supporting role in the ever-evolving relation<br \/>\n  between automatic clients and the human web.\n<\/p>\n<p>\n  The REP &mdash; specifically robots.txt &mdash; became a standard in 2022 as<br \/>\n  <a href=\"https:\/\/datatracker.ietf.org\/doc\/html\/rfc9309\" class=\"external-link\">RFC9309<\/a>.<br \/>\n  However, the heavy lifting was done prior to its standardization: it was the test of time between<br \/>\n  1994 and 2022 that made it popular enough to be adopted by billions of hosts and virtually all<br \/>\n  major crawler operators (excluding adversarial crawlers such as malware scanners). It is a<br \/>\n  straightforward and elegant solution to express preferences with a simple yet versatile syntax.<br \/>\n  In its 25 years of existence it barely had to evolve from its original form, it only got an<br \/>\n  <code>allow<\/code> rule if we only consider the rules that are universally supported by crawlers.\n<\/p>\n<p>\n  That doesn&#8217;t mean that there are no other rules; any crawler operator can come up with their own<br \/>\n  rules. For example, rules like &#8220;<code>clean-param<\/code>&#8221; and &#8220;<code>crawl-delay<\/code>&#8221; are not<br \/>\n  part of RFC9309, but they&#8217;re supported by some search engines &mdash; though not Google Search.<br \/>\n  The &#8220;<code>sitemap<\/code>&#8221; rule, which again is not part of RFC9309, is supported by all major<br \/>\n  search engines. Given enough support, it could become an official rule in the REP.\n<\/p>\n<p>\n  Because the REP can in fact get &#8220;updates&#8221;. It&#8217;s a widely supported protocol and it should grow<br \/>\n  with the internet. Making changes to it is not impossible, but it&#8217;s not easy; it shouldn&#8217;t be<br \/>\n  easy, exactly because the REP is widely supported. Like with any change to a standard, there has<br \/>\n  to be a consensus that changes benefit the majority of the users of the protocol, both on the<br \/>\n  publishers&#8217; and the crawler operators&#8217; side.\n<\/p>\n<p>\n  Due to its simplicity and wide adoption, the REP is an excellent candidate for carrying new<br \/>\n  crawling preferences: billions of publishers are already familiar with robots.txt and its syntax<br \/>\n  for example, so making changes to it comes more naturally for them. On the flip side, crawler<br \/>\n  operators already have robust, well tested parsers and matchers (and Google also open sourced its<br \/>\n  own <a href=\"https:\/\/github.com\/google\/robotstxt\" class=\"external-link\">robots.txt parser<\/a>),<br \/>\n  which means it&#8217;s highly likely that there won&#8217;t be parsing issues with new rules.\n<\/p>\n<p>\n  The same goes for the REP URI level extensions, the <code>X-robots-tag<\/code> HTTP header and its<br \/>\n  meta tag counterpart. If there is a need for a new rule to carry opt-out preferences, they&#8217;re<br \/>\n  easily extensible. How though?\n<\/p>\n<p>\n  The most important thing you, the reader, can do is to talk about your idea publicly and gather<br \/>\n  supporters for that idea. Because the REP is a public standard, no one entity can make unilateral<br \/>\n  changes to it; sure, they can implement support for something new on their side, but that won&#8217;t<br \/>\n  become THE standard. But talking about that change and showing to the ecosystem &mdash; both<br \/>\n  crawler operators and the publishing ecosystem &mdash; that it&#8217;s benefiting everyone will drive<br \/>\n  consensus, and that paves the road to updating the standard.\n<\/p>\n<p>\n  Similarly, if the protocol is lacking something, talk about it publicly. <code>sitemap<\/code><br \/>\n  became a widely supported rule in robots.txt because it was useful for content creators and search<br \/>\n  engines alike, which paved the road to adoption of the extension. If you have a new idea for a<br \/>\n  rule, ask the consumers of robots.txt and creators what they think about it and work with them to<br \/>\n  hash out potential (and likely) issues they raise and write up a proposal.\n<\/p>\n<p>If your driver is to serve the common good, it&#8217;s worth it.<\/p>\n<p class=\"byline-author\">\n  Posted by <a href=\"https:\/\/developers.google.com\/search\/blog\/authors\/gary-illyes\">Gary Illyes<\/a>, Search Relations team\n<\/p>\n<hr class=\"full-width\">\n<h2 id=\"check-out-the-rest-of-the-robots-refresher-series:\" tabindex=\"-1\">Check out the rest of the Robots Refresher series:<\/h2>\n","protected":false},"excerpt":{"rendered":"<p>Friday, March 28, 2025 In the previous posts about the Robots Exclusion Protocol (REP), we explored what you can already do with its various components &mdash; namely robots.txt and the URI level controls. In this post we will explore how the REP can play a supporting role in the ever-evolving relation between automatic clients and [&hellip;]<\/p>\n","protected":false},"menu_order":79825,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-25343","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25343","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25343\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/27463"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=25343"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=25343"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=25343"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}