{"id":25082,"date":"2019-07-02T00:00:00","date_gmt":"2019-07-02T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/a-note-on-unsupported-rules-in-robots-txt\/"},"modified":"2019-07-02T00:00:00","modified_gmt":"2019-07-02T00:00:00","slug":"a-note-on-unsupported-rules-in-robots-txt","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/a-note-on-unsupported-rules-in-robots-txt\/","title":{"rendered":"A note on unsupported rules in robots.txt"},"content":{"rendered":"<p class=\"gargardate\">Tuesday, July 02, 2019<\/p>\n<p>\n   Yesterday we announced that we&#8217;re<br \/>\n   <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/googles-robots-txt-parser-is-now-open-source\/\">open-sourcing Google&#8217;s production robots.txt parser<\/a>.<br \/>\n    It was an exciting moment that paves the road for potential Search open sourcing projects in the<br \/>\n    future! Feedback is helpful, and we&#8217;re eagerly collecting questions from<br \/>\n   <a href=\"https:\/\/github.com\/google\/robotstxt\" class=\"external-link\">developers<\/a> and<br \/>\n   <a href=\"https:\/\/twitter.com\/googlesearchc\" class=\"external-link\">webmasters<\/a> alike. One question<br \/>\n    stood out, which we&#8217;ll address in this post:<br \/>\n   <br \/>\n   Why isn&#8217;t a code handler for other rules like crawl-delay included in the code?\n  <\/p>\n<p>\n   <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/formalizing-the-robots-exclusion-protocol-specification\/\">The internet draft<\/a> we published yesterday provides an<br \/>\n    extensible architecture for rules that are not part of the standard. This means that if a<br \/>\n    crawler wanted to support their own line like <code>unicorns: allowed<\/code>,<br \/>\n    they could. To demonstrate how this would look in a parser, we included a very common line,<br \/>\n    sitemap, in our <a href=\"https:\/\/github.com\/google\/robotstxt\" class=\"external-link\">open-source robots.txt parser<\/a>.\n  <\/p>\n<p>\n   While open-sourcing our parser library, we analyzed the usage of robots.txt rules. In particular,<br \/>\n    we focused on rules unsupported by the internet draft, such as<br \/>\n    <code>crawl-delay<\/code>, <code>nofollow<\/code>, and<br \/>\n    <code>noindex<\/code>. Since these rules were never documented by Google,<br \/>\n    naturally, their usage in relation to Googlebot is very low. Digging further, we saw their usage<br \/>\n    was contradicted by other rules in all but 0.001% of all robots.txt files on the internet.<br \/>\n    These mistakes hurt websites&#8217; presence in Google&#8217;s search results in ways we don&#8217;t think<br \/>\n    webmasters intended.\n  <\/p>\n<p>\n   In the interest of maintaining a healthy ecosystem and preparing for potential future open source<br \/>\n    releases, we&#8217;re retiring all code that handles unsupported and unpublished rules (such as<br \/>\n    <code>noindex<\/code>) on September 1, 2019. For those of you who relied on the<br \/>\n    <code>noindex<\/code> indexing rule in the<br \/>\n    <code>robots.txt<\/code> file, which controls crawling, there are a number of<br \/>\n    alternative options:\n  <\/p>\n<ul>\n<li>\n      <b><a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/block-search-indexing-with-noindex\/\"><code>noindex<\/code><\/a><br \/>\n        in <span>robots<\/span> <code>meta<\/code> tags:<\/b> Supported both in the HTTP response headers and in HTML, the<br \/>\n      <code>noindex<\/code> rule is the most effective way to remove URLs from<br \/>\n      the index when crawling is allowed.\n    <\/li>\n<li>\n      <b><a href=\"https:\/\/en.wikipedia.org\/wiki\/HTTP_404\" class=\"external-link\"><code>404<\/code> and <code>410<\/code> HTTP status codes<\/a>:<\/b><br \/>\n      Both status codes mean that the page does not exist, which will drop such URLs from Google&#8217;s<br \/>\n      index once they&#8217;re crawled and processed.\n    <\/li>\n<li>\n      <b>Password protection:<\/b> Unless markup is used to indicate<br \/>\n      <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/structured-data-for-subscription-and-paywalled-content-creativework\/\">subscription or paywalled content<\/a>,<br \/>\n      hiding a page behind a login will generally remove it from Google&#8217;s index.\n    <\/li>\n<li>\n      <b><code>Disallow<\/code> in <code>robots.txt<\/code>:<\/b> Search engines can only index pages<br \/>\n      that they know about, so blocking the page from being crawled usually means its content won&#8217;t<br \/>\n      be indexed. While the search engine may also index a URL based on links from other pages,<br \/>\n      without seeing the content itself, we aim to make such pages less visible in the future.\n    <\/li>\n<li>\n     <b><a href=\"https:\/\/support.google.com\/webmasters\/answer\/1663419\" class=\"external-link\">Search Console Remove URL tool<\/a>:<\/b><br \/>\n      The tool is a quick and easy method to remove a URL temporarily from Google&#8217;s search results.\n    <\/li>\n<\/ul>\n<p>\n    For more guidance about how to remove information from Google&#8217;s search results, visit our<br \/>\n    <a href=\"https:\/\/developers.google.com\/search\/docs\/guides\/advanced\/remove-information?ref_topic=1724262\" class=\"external-link\">Help Center<\/a>.<br \/>\n    If you have questions, you can find us on <a href=\"https:\/\/twitter.com\/googlesearchc\" class=\"external-link\">Twitter<\/a><br \/>\n    and in our <a href=\"https:\/\/support.google.com\/webmasters\/community\" class=\"external-link\">Webmaster Community<\/a>,<br \/>\n    both <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/about-search-central-live\/\">offline<\/a> and online.\n  <\/p>\n<p class=\"byline-author\">\n    Posted by <a href=\"https:\/\/garyillyes.com\/+\" class=\"external-link\">Gary Illyes<\/a>\n  <\/p>\n","protected":false},"excerpt":{"rendered":"<p>Tuesday, July 02, 2019 Yesterday we announced that we&#8217;re open-sourcing Google&#8217;s production robots.txt parser. It was an exciting moment that paves the road for potential Search open sourcing projects in the future! Feedback is helpful, and we&#8217;re eagerly collecting questions from developers and webmasters alike. One question stood out, which we&#8217;ll address in this post: [&hellip;]<\/p>\n","protected":false},"menu_order":81921,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-25082","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25082","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25082\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/27211"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=25082"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=25082"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=25082"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}