{"id":24283,"date":"2009-12-02T00:00:00","date_gmt":"2009-12-02T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/new-user-agent-for-news\/"},"modified":"2009-12-02T00:00:00","modified_gmt":"2009-12-02T00:00:00","slug":"new-user-agent-for-news","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/new-user-agent-for-news\/","title":{"rendered":"New User Agent for News"},"content":{"rendered":"<aside class=\"key-point\">It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore).<\/aside>\n<p class=\"gargardate\">Wednesday, December 02, 2009<\/p>\n<p>\n  Today we are announcing a new user agent for robots.txt called Googlebot-News that gives<br \/>\n  publishers even more control over their content. In case you haven&#8217;t heard of<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">robots.txt<\/a>, it&#8217;s a web-wide standard that has<br \/>\n  been in use<br \/>\n  <a href=\"https:\/\/en.wikipedia.org\/wiki\/Robots_exclusion_standard\" class=\"external-link\">since 1994<\/a><br \/>\n  and which has support from all major search engines and well-behaved &#8220;robots&#8221; that process the<br \/>\n  web. When a search engine checks whether it has permission to crawl and index a web page, the<br \/>\n  &#8220;check if we&#8217;re allowed to crawl this page&#8221; mechanism is robots.txt.\n<\/p>\n<p>\n  Publishers could easily contact us<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/news_pub\/bin\/answer.py?answer=94003\" class=\"external-link\">via a form<\/a><br \/>\n  if they didn&#8217;t want to be included in Google News but did want to be in Google&#8217;s web search index.<br \/>\n  Now, publishers can manage their content in Google News in an even more automated way. Site owners<br \/>\n  can just add <code>Googlebot-News<\/code> specific rules to their robots.txt file. Similar to<br \/>\n  the <code>Googlebot<\/code> and <code>Googlebot-Image<\/code> user agents, the new<br \/>\n  <code>Googlebot-News<\/code> user agent can be used to specify which pages of a website should be<br \/>\n  crawled and ultimately appear in Google News.\n<\/p>\n<p>Here are a few examples for publishers:<\/p>\n<p><b>Include pages in both Google web search and News:<\/b><\/p>\n<div><\/div>\n<p>This is the easiest case. In fact, a robots.txt file is not even required for this case.<\/p>\n<p><b>Include pages in Google web search, but not in News:<\/b><\/p>\n<div><\/div>\n<p>\n  This robots.txt file says that no files are disallowed from Google&#8217;s general web crawler,<br \/>\n  called <code>Googlebot<\/code>, but the user agent <code>Googlebot-News<\/code> is blocked from all<br \/>\n  files on the website.\n<\/p>\n<p><b>Include pages in Google News, but not Google web search:<\/b><\/p>\n<div><\/div>\n<p>\n  When parsing a robots.txt file, Google obeys the most specific rule. The first two lines<br \/>\n  tell us that Googlebot (the user agent for Google&#8217;s web index) is blocked from crawling any pages<br \/>\n  from the site. The next rule, which applies to the more specific user agent for Google News,<br \/>\n  overrides the blocking of Googlebot and gives permission for Google News to crawl pages from the<br \/>\n  website.\n<\/p>\n<p><b>Block different sets of pages from Google web search and Google News:<\/b><\/p>\n<div><\/div>\n<p>\n  The pages blocked from Google web search and Google News can be controlled independently. This<br \/>\n  robots.txt file blocks recent news articles (URLs in the \/latest_news folder) from Google web<br \/>\n  search, but allows them to appear on Google News. Conversely, it blocks premium content (URLs in<br \/>\n  the \/archives folder) from Google News, but allows them to appear in Google web search.\n<\/p>\n<p><b>Stop Google web search and Google News from crawling pages:<\/b><\/p>\n<div><\/div>\n<p>\n  This robots.txt file tells Google that Googlebot, the user agent for our web search crawler,<br \/>\n  should not crawl any pages from the site. Because no specific rule for Googlebot-News is<br \/>\n  given, our News search will abide by the general guidance for Googlebot and will not crawl pages<br \/>\n  for Google News.\n<\/p>\n<p>\n  For some queries, we display results from Google News in a discrete box or section on the web<br \/>\n  search results page, along with our regular web search results. We sometimes do this for Images,<br \/>\n  Videos, Maps, and Products, too. This is known as<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/webmasters\/bin\/answer.py?answer=159206\" class=\"external-link\">Universal search results<\/a>.<br \/>\n  Since Google News powers Universal &#8220;News&#8221; search results, if you block the<br \/>\n  <code>Googlebot-News<\/code> user agent then your site&#8217;s news stories won&#8217;t be included in<br \/>\n  Universal search results.\n<\/p>\n<p>\n  We are currently testing our support for the new user agent. If you see any problems<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/community\" class=\"external-link\">please let us know<\/a>.<br \/>\n  Note that<br \/>\n  <a href=\"https:\/\/www.youtube.com\/watch?v=KBdEwpRQRD0\" class=\"external-link\">it is possible for Google<\/a><br \/>\n  to return a link to a page in some situations even when we didn&#8217;t crawl that page. If you&#8217;d like<br \/>\n  to<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">read more about robots.txt<\/a>, we provide<br \/>\n  additional documentation on our website. We hope webmasters will enjoy the flexibility and easier<br \/>\n  management that the Googlebot-News user agent provides.\n<\/p>\n<p class=\"byline-author\">\n  Written by<br \/>\n  <a href=\"https:\/\/developers.google.com\/search\/blog\/authors\/jonathan-simon\" rel=\"author\" class=\"external-link\">Jonathan Simon<\/a>,<br \/>\n  Webmaster Trends Analyst<\/p>\n","protected":false},"excerpt":{"rendered":"<p>It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore). Wednesday, December 02, 2009 Today we are announcing a new user agent for robots.txt called Googlebot-News that gives publishers even more control over their [&hellip;]<\/p>\n","protected":false},"menu_order":85420,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-24283","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24283","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24283\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26758"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=24283"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=24283"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=24283"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}