{"id":24000,"date":"2008-06-03T00:00:00","date_gmt":"2008-06-03T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/improving-on-robots-exclusion-protocol\/"},"modified":"2008-06-03T00:00:00","modified_gmt":"2008-06-03T00:00:00","slug":"improving-on-robots-exclusion-protocol","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/improving-on-robots-exclusion-protocol\/","title":{"rendered":"Improving on Robots Exclusion Protocol"},"content":{"rendered":"<aside class=\"key-point\">It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore).<br \/>\nRead the up-to-date documentation about <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">robots.txt<\/a>.<\/aside>\n<p class=\"gargardate\">Tuesday, June 03, 2008<\/p>\n<p>\n  Web publishers often ask us how they can maximize their visibility on the web. Much of this has<br \/>\n  to do with search engine optimization&mdash;making sure a publisher&#8217;s content shows up on all the<br \/>\n  search engines.\n<\/p>\n<p>\n  However, there are some cases in which publishers need to communicate more information to search<br \/>\n  engines&mdash;like the fact that they <i>don&#8217;t<\/i> want certain content to appear in search<br \/>\n  results. And for that they use something called the<br \/>\n  <a href=\"https:\/\/googleblog.blogspot.com\/2007\/01\/controlling-how-search-engines-access.html\" class=\"external-link\">Robots Exclusion Protocol (REP)<\/a>,<br \/>\n  which lets publishers control how search engines access their site: whether it&#8217;s controlling the<br \/>\n  visibility of their content across their site (via robots.txt) or down to a much more granular<br \/>\n  level for individual pages (via <code>meta<\/code> tags).\n<\/p>\n<p>\n  Since it was introduced in the early &#8217;90s, REP has become the de facto standard by which web<br \/>\n  publishers specify which parts of their site they want public and which parts they want to keep<br \/>\n  private. Today, millions of publishers use REP as an easy and efficient way to communicate with<br \/>\n  search engines. Its strength lies in its flexibility to evolve in parallel with the web, its<br \/>\n  universal implementation across major search engines and all major robots, and in the way it<br \/>\n  works for any publisher, no matter how large or small.\n<\/p>\n<p>\n  While REP is observed by virtually all search engines, we&#8217;ve never come together to detail how<br \/>\n  we each interpret different tags. Over the last couple of years, we have worked with Microsoft<br \/>\n  and Yahoo! to bring forward standards such as<br \/>\n  <a href=\"https:\/\/www.sitemaps.org\/\" class=\"external-link\">Sitemaps<\/a><br \/>\n  and offer additional tools for webmasters. Since the original announcement, we have, and will<br \/>\n  continue to, deliver further improvements based on what we are hearing from the community.\n<\/p>\n<p>\n  Today, in that same spirit of making the lives of webmasters simpler, we&#8217;re releasing detailed<br \/>\n  documentation about how we implement REP. This will provide a common implementation for<br \/>\n  webmasters and make it easier for any publisher to know how their REP rules will be handled<br \/>\n  by three major search providers&mdash;making REP more intuitive and friendly to even more<br \/>\n  publishers on the web.\n<\/p>\n<p>So, without further ado&#8230;<\/p>\n<h2 id=\"common-rep-rules\" tabindex=\"-1\">Common REP rules<\/h2>\n<p>\n  The following list are all the major REP features currently implemented by Google, Microsoft,<br \/>\n  and Yahoo!.  With each feature, you&#8217;ll see what it does and how you should communicate it.\n<\/p>\n<p>\n  Each of these rules can be specified to be applicable for all crawlers or for specific<br \/>\n  crawlers by targeting them to specific user-agents, which is how any crawler identifies itself.<br \/>\n  Apart from the identification by user-agent, each of our crawlers also supports Reverse DNS based<br \/>\n  authentication to allow you to verify the identity of the crawler.\n<\/p>\n<h3 id=\"robots.txt-rules\" tabindex=\"-1\">Robots.txt rules<\/h3>\n<table>\n<tbody>\n<tr>\n<td width=\"25%\"><b>Rule<\/b><\/td>\n<td width=\"37%\"><b>Impact<\/b><\/td>\n<td width=\"37%\"><b>Use cases<\/b><\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\"><code>Disallow<\/code><\/a>\n      <\/td>\n<td width=\"37%\">\n        Tells a crawler not to index your site&mdash;your site&#8217;s robots.txt file still needs to be<br \/>\n        crawled to find this rule, however disallowed pages will not be crawled\n      <\/td>\n<td width=\"37%\">\n        &#8216;No Crawl&#8217; page from a site. This rule in the default syntax prevents specific path(s)<br \/>\n        of a site from being crawled.\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/google-crawlers\/\"><code>Allow<\/code><\/a>\n      <\/td>\n<td width=\"37%\">\n        Tells a crawler the specific pages on your site you want indexed so you can use this in<br \/>\n        combination with Disallow\n      <\/td>\n<td width=\"37%\">\n        This is useful in particular in conjunction with Disallow clauses, where a large section of<br \/>\n        a site is disallowed except for a small section within it\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-syntax\/#url-matching-based-on-path-values\"><code>$<\/code> Wildcard Support<\/a>\n      <\/td>\n<td width=\"37%\">\n        Tells a crawler to match everything from the end of a URL&mdash;large number of directories<br \/>\n        without specifying specific pages\n      <\/td>\n<td width=\"37%\">\n        &#8216;No Crawl&#8217; files with specific patterns, for example, files with certain filetypes that<br \/>\n        always have a certain extension, say pdf\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-syntax\/#url-matching-based-on-path-values\"><code>*<\/code> Wildcard Support<\/a>\n      <\/td>\n<td width=\"37%\">\n        Tells a crawler to match a sequence of characters\n      <\/td>\n<td width=\"37%\">\n        &#8216;No Crawl&#8217; URLs with certain patterns, for example, disallow URLs with session ids or other<br \/>\n        extraneous parameters\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/sitemaps-overview\/\"><code>Sitemaps<\/code> Standort<\/a>\n      <\/td>\n<td width=\"37%\">Tells a crawler where it can find your Sitemaps<\/td>\n<td width=\"37%\">\n        Point to other locations where feeds exist to help crawlers find URLs on a site\n      <\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h3 id=\"html-meta-rules\" tabindex=\"-1\">HTML <code>meta<\/code> rules<\/h3>\n<table>\n<tbody>\n<tr>\n<td width=\"25%\"><b>Rule<\/b><\/td>\n<td width=\"37%\"><b>Impact<\/b><\/td>\n<td width=\"37%\"><b>Use cases<\/b><\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/block-search-indexing-with-noindex\/\"><code>noindex<\/code> <code>meta<\/code> tag<\/a>\n      <\/td>\n<td width=\"37%\">Tells a crawler not to index a given page<\/td>\n<td width=\"37%\">\n        Don&#8217;t index the page. This allows pages that are crawled to be kept out of the index.\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/rel-attributes\/\"><code>nofollow<\/code> <code>meta<\/code> tag<\/a>\n      <\/td>\n<td width=\"37%\">Tells a crawler not to follow a link to other content on a given page<\/td>\n<td>\n        Prevent publicly writeable areas to be abused by spammers looking for link credit. By using<br \/>\n        <code>nofollow<\/code> you let the robot know that you are discounting all outgoing links<br \/>\n        from this page.\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-meta-tag\/#nosnippet\"><code>nosnippet<\/code> <code>meta<\/code> tag<\/a>\n      <\/td>\n<td width=\"37%\">\n        Tells a crawler not to display snippets in the search results for a given page\n      <\/td>\n<td width=\"37%\">Present no snippet for the page on Search Results<\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-meta-tag\/#noarchive\"><code>noarchive<\/code> <code>meta<\/code> tag<\/a>\n      <\/td>\n<td width=\"37%\">Tells a search engine not to show a &#8220;cached&#8221; link for a given page<\/td>\n<td width=\"37%\">\n        Do not make available to users a copy of the page from the Search Engine cache\n      <\/td>\n<\/tr>\n<tr>\n<td width=\"25%\">\n        <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-meta-tag\/#preventdmoz\"><code>noodp<\/code> <code>meta<\/code> tag<\/a>\n      <\/td>\n<td width=\"37%\">\n        Tells a crawler not to use a title and snippet from the Open Directory Project for a given<br \/>\n        page\n      <\/td>\n<td width=\"37%\">\n        Do not use the ODP (Open Directory Project) title and snippet for this page\n      <\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>\n  These rules are applicable for all forms of content. They can be placed in either the HTML<br \/>\n  of a page or in the HTTP header for non-HTML content, for example, PDF, video, etc. using an<br \/>\n  <code>X-Robots-Tag<\/code>. You can read more about it here:<br \/>\n  <a href=\"https:\/\/googleblog.blogspot.com\/2007\/07\/robots-exclusion-protocol-now-with-even.html\" class=\"external-link\"><code>X-Robots-Tag<\/code> Post<\/a><br \/>\n  or in<br \/>\n  <a href=\"https:\/\/googleblog.blogspot.com\/2007\/02\/robots-exclusion-protocol.html\" class=\"external-link\">our series of posts<\/a><br \/>\n  about using robots and <code>meta<\/code> tags.\n<\/p>\n<h2 id=\"other-rep-rules\" tabindex=\"-1\">Other REP rules<\/h2>\n<p>\n  The rules listed above are used by Microsoft, Google, and Yahoo!, but may not be implemented<br \/>\n  by all other search engines. In addition, the following rules are supported by Google, but<br \/>\n  are not supported by all three as are those above:\n<\/p>\n<p>\n  <b><code>unavailable_after<\/code> <code>meta<\/code> tag<\/b> &#8211; Tells a crawler<br \/>\n  <a href=\"https:\/\/googleblog.blogspot.com\/2007\/07\/robots-exclusion-protocol-now-with-even,html\" class=\"external-link\">when a page should &#8220;expire&#8221;<\/a>,<br \/>\n  that is, after which date it should not show up in search results.\n<\/p>\n<p>\n  <b><code>noimageindex<\/code> <code>meta<\/code> tag<\/b> &#8211; Tells a crawler<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-meta-tag\/#noimageindex\">not to index images for a given page<\/a><br \/>\n  in search results.\n<\/p>\n<p>\n  <b><code>notranslate<\/code> <code>meta<\/code> tag<\/b> &#8211; Tells a crawler<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-meta-tag\/#notranslate\">not to translate the content on a page into different languages<\/a><br \/>\n  for search results.\n<\/p>\n<p>\n  Going forward, we plan to continue to work together to ensure that as new uses of REP arise, we&#8217;re<br \/>\n  able to make it as easy as possible for webmasters to use them.  So stay tuned for more!\n<\/p>\n<h2 id=\"learn-more\" tabindex=\"-1\">Learn more<\/h2>\n<p>\n  You can find out more about robots.txt in our<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/overview-of-crawling-and-indexing-topics\/\">documentation<\/a> and at<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/how-search-works\/#crawling\">Google&#8217;s Webmaster help center<\/a>,<br \/>\n  which contains lots of helpful information, including:\n<\/p>\n<ul>\n<li>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/how-to-write-and-submit-a-robots-txt-file\/\">How to create a robots.txt file<\/a>\n  <\/li>\n<li>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/google-crawlers\/\">Descriptions of each user-agent that Google uses<\/a>\n  <\/li>\n<li>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-syntax\/#url-matching-based-on-path-values\">How to use pattern matching<\/a>\n  <\/li>\n<li>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/update-your-robots-txt-file\/#refresh-googles-robots.txt-cache\">How often we recrawl your robots.txt file<\/a>\n  <\/li>\n<\/ul>\n<p>\n  We&#8217;ve also done several posts in our <a href=\"https:\/\/developers.google.com\/search\/blog\">webmaster blog<\/a> about robots.txt<br \/>\n  that you may find useful, such as:\n<\/p>\n<ul>\n<li><a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/using-a-robots-txt-file\/\">Using robots.txt files<\/a><\/li>\n<li><a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/all-about-googlebot\/\">All about Googlebot<\/a><\/li>\n<\/ul>\n<p>\n  There is also a<br \/>\n  <a href=\"https:\/\/www.robotstxt.org\/db\" class=\"external-link\">useful list of the bots<\/a><br \/>\n  used by the major search engines.\n<\/p>\n<p>\n  To see what our colleagues have to say, you can also check out the blog posts published by<br \/>\n  <a href=\"https:\/\/www.ysearchblog.com\/archives\/000587\" class=\"external-link\">Yahoo!<\/a> and<br \/>\n  <a href=\"https:\/\/blogs.msdn.com\/webmaster\/archive\/2008\/06\/03\/robots-exclusion-protocol-joining-together-to-provide-better-documentation.aspx\" class=\"external-link\">Microsoft<\/a>.\n<\/p>\n<p class=\"byline-author\">Written by Prashanth Koppula, Product Manager<\/p>\n","protected":false},"excerpt":{"rendered":"<p>It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore). Read the up-to-date documentation about robots.txt. Tuesday, June 03, 2008 Web publishers often ask us how they can maximize their visibility on the web. [&hellip;]<\/p>\n","protected":false},"menu_order":85967,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-24000","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24000","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24000\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26608"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=24000"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=24000"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=24000"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}