{"id":25085,"date":"2019-07-01T00:00:00","date_gmt":"2019-07-01T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/googles-robots-txt-parser-is-now-open-source\/"},"modified":"2019-07-01T00:00:00","modified_gmt":"2019-07-01T00:00:00","slug":"googles-robots-txt-parser-is-now-open-source","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/googles-robots-txt-parser-is-now-open-source\/","title":{"rendered":"Google&#8217;s robots.txt parser is now open source"},"content":{"rendered":"<p class=\"gargardate\">Monday, July 01, 2019<\/p>\n<p>\n  For 25 years, the <a href=\"https:\/\/www.robotstxt.org\/norobots-rfc.txt\" class=\"external-link\">Robots Exclusion Protocol (REP)<\/a><br \/>\n  was only a de-facto standard. This had frustrating implications sometimes. On one hand, for<br \/>\n  webmasters, it meant uncertainty in corner cases, like when their text editor included<br \/>\n  <a href=\"https:\/\/en.wikipedia.org\/wiki\/Byte_order_mark\" class=\"external-link\">BOM<\/a> characters in<br \/>\n  their robots.txt files. On the other hand, for crawler and tool developers, it also brought<br \/>\n  uncertainty; for example, how should they deal with robots.txt files that are hundreds of<br \/>\n  megabytes large?\n<\/p>\n<p><img decoding=\"async\" alt=\"Googlebot unboxing a website\" height=\"320\"\n     src=\"https:\/\/alienroad.com\/wp-content\/uploads\/kb-gorsel\/g-6bb39cb08b83.png\" loading=\"lazy\" width=\"640\"\/><\/p>\n<p>\n  Today, <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/formalizing-the-robots-exclusion-protocol-specification\/\">we announced<\/a> that we&#8217;re spearheading the effort<br \/>\n  to make the REP an internet standard. While this is an important step, it means extra work for<br \/>\n  developers who parse robots.txt files.\n<\/p>\n<p>\n  We&#8217;re here to help: we <a href=\"https:\/\/github.com\/google\/robotstxt\" class=\"external-link\">open sourced<\/a><br \/>\n  the C++ library that our production systems use for parsing and matching rules in robots.txt<br \/>\n  files. This library has been around for 20 years and it contains pieces of code that were written<br \/>\n  in the 90&#8217;s. Since then, the library evolved; we learned a lot about how webmasters write<br \/>\n  robots.txt files and corner cases that we had to cover for, and added what we learned over the<br \/>\n  years also to the internet draft when it made sense.\n<\/p>\n<p>\n  We also included a testing tool in the open source package to help you test a few rules. Once<br \/>\n  built, the usage is very straightforward:\n<\/p>\n<p>\n<code>robots_main &lt;robots.txt content&gt; &lt;user_agent&gt; &lt;url&gt;<\/code>\n<\/p>\n<p>\n  If you want to check out the library, head over to our GitHub repository for the<br \/>\n  <a href=\"https:\/\/github.com\/google\/robotstxt\" class=\"external-link\">robots.txt parser<\/a>. We&#8217;d love<br \/>\n  to see what you can build using it! If you built something using the library, drop us a comment on<br \/>\n  <a href=\"https:\/\/twitter.com\/googlesearchc\" class=\"external-link\">Twitter<\/a>, and if you have comments<br \/>\n  or questions about the library, find us on<br \/>\n  <a href=\"https:\/\/github.com\/google\/robotstxt\" class=\"external-link\">GitHub<\/a>.\n<\/p>\n<p class=\"byline-author\">\n  Posted by <a href=\"https:\/\/twitter.com\/epere4\" class=\"external-link\">Edu Pereda<\/a>,<br \/>\n  <a href=\"https:\/\/github.com\/lvandeve\" class=\"external-link\">Lode Vandevenne<\/a>, and<br \/>\n  <a href=\"https:\/\/garyillyes.com\/+\" class=\"external-link\">Gary Illyes<\/a>, Search Open Sourcing team<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Monday, July 01, 2019 For 25 years, the Robots Exclusion Protocol (REP) was only a de-facto standard. This had frustrating implications sometimes. On one hand, for webmasters, it meant uncertainty in corner cases, like when their text editor included BOM characters in their robots.txt files. On the other hand, for crawler and tool developers, it [&hellip;]<\/p>\n","protected":false},"menu_order":81922,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-25085","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25085","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/25085\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/27214"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=25085"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=25085"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=25085"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}