{"id":23962,"date":"2008-03-01T00:00:00","date_gmt":"2008-03-01T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/how-to-use-robots-txt\/"},"modified":"2008-03-01T00:00:00","modified_gmt":"2008-03-01T00:00:00","slug":"how-to-use-robots-txt","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/how-to-use-robots-txt\/","title":{"rendered":"How to use robots.txt"},"content":{"rendered":"<p class=\"gargardate\">Saturday, March 1, 2008<\/p>\n<p>\n  A <a href=\"https:\/\/www.robotstxt.org\/\" class=\"external-link\">robots.txt<\/a><br \/>\n  provides restrictions to search engine robots (known as &#8220;bots&#8221;) that crawl the web. These bots are<br \/>\n  automated, and before they access pages of a site, they check to see if a robots.txt file exists<br \/>\n  that prevents them from accessing certain pages. If you want to protect some of your contents<br \/>\n  from being indexed by search engines, robots.txt is a simple tool for it. In this time, we would<br \/>\n  like to discuss how to use it.\n<\/p>\n<h2 id=\"placing-robots.txt\" tabindex=\"-1\">Placing Robots.txt<\/h2>\n<p>\n  The robots.txt file is a text file, with one or more records. The robots.txt file must be reside<br \/>\n  in the root of the domain and must be exactly named robots.txt. A robots.txt file located in a<br \/>\n  subdirectory is not a valid, as bots only check for this file in the root of the domain.\n<\/p>\n<p>\n  For instance,<br \/>\n  <code>https:\/\/www.example.com\/robots.txt<\/code> is a valid location. But,<br \/>\n  <code>https:\/\/www.example.com\/mysite\/robots.txt<\/code> is not.\n<\/p>\n<p>Example of a robots.txt:<\/p>\n<div><\/div>\n<h2 id=\"block-or-remove-your-entire-website-using-a-robots.txt-file\" tabindex=\"-1\">Block or remove your entire website using a robots.txt file<\/h2>\n<p>\n  To remove your site from search engines and prevent all robots from crawling it in the future,<br \/>\n  place the following robots.txt file in your server root:\n<\/p>\n<div><\/div>\n<p>\n  To remove your site from Google only and prevent just Googlebot from crawling your site in the<br \/>\n  future, place the following robots.txt file in your server root:\n<\/p>\n<div><\/div>\n<p>\n  Each port must have its own robots.txt file. In particular, if you serve content via both http<br \/>\n  and https, you&#8217;ll need a separate robots.txt file for each of these protocols. For example, to<br \/>\n  allow Googlebot to index all http pages but no https pages, you&#8217;d use the robots.txt files below.\n<\/p>\n<p>For your http protocol (<code>https:\/\/example.com\/robots.txt<\/code>):<\/p>\n<div><\/div>\n<p>For the https protocol (<code>https:\/\/yourserver.com\/robots.txt<\/code>):<\/p>\n<div><\/div>\n<p>Allow all robots complete access<\/p>\n<div><\/div>\n<p>(alternative solution: Just create an empty &#8220;\/robots.txt&#8221; file, or don&#8217;t use one at all.)<\/p>\n<h2 id=\"block-or-remove-pages-using-a-robots.txt-file\" tabindex=\"-1\">Block or remove pages using a robots.txt file<\/h2>\n<p>You can use a robots.txt file to block Googlebot from crawling pages on your site.<\/p>\n<p>\n  For example, if you&#8217;re manually creating a robots.txt file, to block Googlebot from crawling all<br \/>\n  pages under a particular directory (for example, private), you&#8217;d use the following robots.txt<br \/>\n  entry:\n<\/p>\n<div><\/div>\n<p>\n  To block Googlebot from crawling all files of a specific file type (for example, .gif), you&#8217;d use<br \/>\n  the following robots.txt entry:\n<\/p>\n<div><\/div>\n<p>\n  To block Googlebot from crawling any URL that includes a <code>?<\/code> (more specifically, any<br \/>\n  URL that begins with your domain name, followed by any string, followed by a question mark,<br \/>\n  followed by any string):\n<\/p>\n<div><\/div>\n<p>\n  While we won&#8217;t crawl or index the content of pages blocked by robots.txt, we may still crawl and<br \/>\n  index the URLs if we find them on other pages on the web. As a result, the URL of the page or<br \/>\n  other publicly available information such as anchor text in links to the site can appear in<br \/>\n  Google search results. However, no content from your pages will be crawled, indexed, or displayed.\n<\/p>\n<p>\n  As a part of webmaster tool, Google provides<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/webmasters\/bin\/answer.py&amp;answer=35237\" class=\"external-link\">robots.txt analysis tool<\/a>.<br \/>\n  The tool reads the robots.txt file in the same way Googlebot does and gives you results for<br \/>\n  Google user-agents. We strongly suggest to use it. Before creating robots.txt, you should think<br \/>\n  about how much information you want to share with people, or to keep private. Remember that search<br \/>\n  engine is a good way to have your contents publicly more accessible. By using robots.txt properly,<br \/>\n  people will be happy to visit your website through search engine but meanwhile you can still<br \/>\n  prevent your private information from being exposed.\n<\/p>\n<p class=\"byline-author\">By Chao Ma, In Hyuk Seok<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Saturday, March 1, 2008 A robots.txt provides restrictions to search engine robots (known as &#8220;bots&#8221;) that crawl the web. These bots are automated, and before they access pages of a site, they check to see if a robots.txt file exists that prevents them from accessing certain pages. If you want to protect some of your [&hellip;]<\/p>\n","protected":false},"menu_order":86061,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23962","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23962","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23962\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26580"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23962"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23962"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23962"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}