{"id":23783,"date":"2006-02-24T00:00:00","date_gmt":"2006-02-24T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/using-a-robots-txt-file\/"},"modified":"2006-02-24T00:00:00","modified_gmt":"2006-02-24T00:00:00","slug":"using-a-robots-txt-file","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/using-a-robots-txt-file\/","title":{"rendered":"Using a robots.txt file"},"content":{"rendered":"<p class=\"gargardate\">February 24, 2006<\/p>\n<p>\n  A couple of weeks ago, we<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/more-stats-and-analysis-of-robots-txt-files\/\">launched a robots.txt<\/a><br \/>\n  analysis tool. This tool gives you information about how<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/googlebot\/\">Googlebot<\/a><br \/>\n  interprets your robots.txt file. You can read more about the<br \/>\n  <a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9309.html\" class=\"external-link\">robots.txt Robots Exclusion Standard<\/a>,<br \/>\n  but we thought we&#8217;d answer some common questions here.\n<\/p>\n<h2 id=\"what-is-a-robots.txt-file\" tabindex=\"-1\">What is a robots.txt file?<\/h2>\n<p>\n  A robots.txt file provides restrictions to search engine robots (known as &#8220;bots&#8221;) that crawl the<br \/>\n  web. These bots are automated, and before they access pages of a site, they check to see if a<br \/>\n  robots.txt file exists that prevents them from accessing certain pages.\n<\/p>\n<h2 id=\"does-my-site-need-a-robots.txt-file\" tabindex=\"-1\">Does my site need a robots.txt file?<\/h2>\n<p>\n  Only if your site includes content that you don&#8217;t want search engines to index. If you want<br \/>\n  search engines to index everything in your site, you don&#8217;t need a robots.txt file (not even an<br \/>\n  empty one).\n<\/p>\n<h2 id=\"where-should-the-robots.txt-file-be-located\" tabindex=\"-1\">Where should the robots.txt file be located?<\/h2>\n<p>\n  The robots.txt file <em>must<\/em> reside in the root of the domain. A robots.txt file located in<br \/>\n  a subdirectory isn&#8217;t valid, as bots only check for this file in the root of the domain. For<br \/>\n  instance, <code>https:\/\/www.example.com\/robots.txt<\/code> is a valid location. But,<br \/>\n  <code>https:\/\/www.example.com\/mysite\/robots.txt<\/code> is not. If you don&#8217;t have access to the<br \/>\n  root of a domain, you can restrict access using the<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-meta-tag\/\">Robots META tag.<\/a>\n<\/p>\n<h2 id=\"how-do-i-create-a-robots.txt-file\" tabindex=\"-1\">How do I create a robots.txt file?<\/h2>\n<p>\n  You can create this file in any text editor. It should be an<br \/>\n  <a href=\"https:\/\/www.google.com\/search?q=ascii+text+file\" class=\"external-link\">ASCII-encoded text file<\/a>,<br \/>\n  not an HTML file. The filename should be lowercase.\n<\/p>\n<h2 id=\"what-should-the-syntax-of-my-robots.txt-file-be\" tabindex=\"-1\">What should the syntax of my robots.txt file be?<\/h2>\n<p>The simplest robots.txt file uses two rules:<\/p>\n<ul>\n<li><strong><code>User-Agent:<\/code><\/strong> the robot the following rule applies to<\/li>\n<li><strong><code>Disallow:<\/code><\/strong> the pages you want to block<\/li>\n<\/ul>\n<p>\n  These two lines are considered a single entry in the file. You can include as many entries as you<br \/>\n  want. You can include multiple Disallow lines in one entry.\n<\/p>\n<h2 id=\"user-agent\" tabindex=\"-1\"><code>User-Agent<\/code><\/h2>\n<p>\n  A user-agent is a specific search engine robot. The<br \/>\n  <a href=\"https:\/\/www.robotstxt.org\/wc\/active.html\" class=\"external-link\">Web Robots Database<\/a><br \/>\n  lists many common bots. You can set an entry to apply to a specific bot (by listing the name) or<br \/>\n  you can set it to apply to all bots (by listing an asterisk). An entry that applies to all bots<br \/>\n  looks like this:\n<\/p>\n<div><\/div>\n<h2 id=\"disallow\" tabindex=\"-1\"><code>Disallow<\/code><\/h2>\n<p>\n  The <code>Disallow<\/code> line lists the pages you want to block. You can list a specific URL or<br \/>\n  a pattern. The entry should begin with a forward slash (\/).\n<\/p>\n<ul>\n<li>\n    <strong>To block the entire site<\/strong>, use a forward slash: <code>Disallow: \/<\/code>\n  <\/li>\n<li>\n    <strong>To block a directory<\/strong>, follow the directory name with a forward slash:<br \/>\n    <code>Disallow: \/private_directory\/<\/code>\n  <\/li>\n<li>\n    <strong>To block a page<\/strong>, list the page:<br \/>\n    <code>Disallow: \/private_file.html<\/code>\n  <\/li>\n<\/ul>\n<p>\n  URLs are case-sensitive. For instance, <code>Disallow: \/private_file.html<\/code> would block<br \/>\n  <code>https:\/\/www.example.com\/private_file.html<\/code>, but would allow<br \/>\n  <code>https:\/\/www.example.com\/Private_File.html<\/code>.\n<\/p>\n<h2 id=\"how-do-i-block-googlebot\" tabindex=\"-1\">How do I block Googlebot?<\/h2>\n<p>\n  Google uses several user-agents. You can block access to any of them by including the bot name on<br \/>\n  the User-Agent line of an entry.\n<\/p>\n<ul>\n<li>\n    <strong>Googlebot:<\/strong> crawl pages from our<br \/>\n    <a href=\"https:\/\/www.google.com\/\" class=\"external-link\">web index<\/a>.\n  <\/li>\n<li>\n    <strong>Googlebot-Mobile:<\/strong> crawls pages for our<br \/>\n    <a href=\"https:\/\/mobile.google.com\/mobile_search.html\" class=\"external-link\">mobile index<\/a>.\n  <\/li>\n<li>\n    <strong>Googlebot-Image:<\/strong> crawls pages for our<br \/>\n    <a href=\"https:\/\/www.google.com\/imghp?tab=wi&#038;q=\" class=\"external-link\">image index<\/a>.\n  <\/li>\n<li>\n    <strong>Mediapartners-Google:<\/strong> crawls pages to determine<br \/>\n    <a href=\"https:\/\/www.google.com\/adsense\/\" class=\"external-link\">AdSense content<\/a><br \/>\n    (used only if you show AdSense ads on your site).\n  <\/li>\n<\/ul>\n<h2 id=\"can-i-allow-pages\" tabindex=\"-1\">Can I allow pages?<\/h2>\n<p>\n  Yes, Googlebot recognizes an extension to the robots.txt standard called <code>Allow<\/code>. This<br \/>\n  extension may not be recognized by all other search engine bots, so check with other search<br \/>\n  engines you&#8217;re interested in to find out. The <code>Allow<\/code> line works exactly like the<br \/>\n  <code>Disallow<\/code> line. Simply list a directory or page you want to allow.\n<\/p>\n<p>\n  You may want to use <code>Disallow<\/code> and <code>Allow<\/code> together. For instance, to block<br \/>\n  access to all pages in a subdirectory except one, you could use the following entries:\n<\/p>\n<div><\/div>\n<p>\n  Those entries would block all pages inside the folder1 directory except for<br \/>\n  <code>myfile.html<\/code>.\n<\/p>\n<h2 id=\"i-dont-want-certain-pages-of-my-site-to-be-indexed,-but-i-want-to-show-adsense-ads-on-those-pages.-can-i-do-that\" tabindex=\"-1\">\n  I don&#8217;t want certain pages of my site to be indexed, but I want to show AdSense ads on those<br \/>\n  pages. Can I do that?<br \/>\n<\/h2>\n<p>\n  Yes, you can <code>Disallow<\/code> all bots other than <code>Mediapartners-Google<\/code> from<br \/>\n  those pages. This keeps the pages from being indexed, but lets<br \/>\n  <code>Googlebot-MediaPartners<\/code> bot analyze the pages to determine the ads to show.<br \/>\n  <code>Googlebot-MediaPartners<\/code> bot doesn&#8217;t share pages with the other Google user-agents.<br \/>\n  For instance, you could use the following entries:\n<\/p>\n<div><\/div>\n<h2 id=\"i-dont-want-to-list-every-file-that-i-want-to-block.-can-i-use-pattern-matching\" tabindex=\"-1\">I don&#8217;t want to list every file that I want to block. Can I use pattern matching?<\/h2>\n<p>\n  Yes, Googlebot interprets some pattern matching. This is an extension of the standard, so not all<br \/>\n  bots may follow it.\n<\/p>\n<h2 id=\"matching-a-sequence-of-characters\" tabindex=\"-1\">Matching a sequence of characters<\/h2>\n<p>\n  You can use an asterisk (<code>*<\/code>) to match a sequence of characters. For instance, to block<br \/>\n  access to all subdirectories that begin with <code>private<\/code>, you could use the following<br \/>\n  entry:\n<\/p>\n<div><\/div>\n<h2 id=\"how-can-i-make-sure-that-my-file-blocks-and-allows-what-i-want-it-to\" tabindex=\"-1\">How can I make sure that my file blocks and allows what I want it to?<\/h2>\n<p>\n  You can use our<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/analyzing-a-robots-txt-file\/\">robots.txt analysis tool<\/a> to:\n<\/p>\n<ul>\n<li>Check specific URLs to see if your robots.txt file allows or blocks them.<\/li>\n<li>See if Googlebot had trouble parsing any lines in your robots.txt file.<\/li>\n<li>Test changes to your robots.txt file.<\/li>\n<\/ul>\n<p>\n  Also, if you don&#8217;t currently use a robots.txt file, you can create one and then test it with the<br \/>\n  tool before you upload it to your site.\n<\/p>\n<h2 id=\"if-i-change-my-robots.txt-file-or-upload-a-new-one,-how-soon-will-it-take-effect\" tabindex=\"-1\">If I change my robots.txt file or upload a new one, how soon will it take effect?<\/h2>\n<p>\n  We generally download robots.txt files about once a day. You can see the last time we downloaded<br \/>\n  your file by accessing the robots.txt tab in your Sitemaps account and checking the<br \/>\n  <strong>Last downloaded<\/strong> date and time.\n<\/p>\n<p class=\"byline-author\">Posted by <a href=\"https:\/\/www.vanessafox.com\/\" class=\"external-link\">Vanessa Fox<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>February 24, 2006 A couple of weeks ago, we launched a robots.txt analysis tool. This tool gives you information about how Googlebot interprets your robots.txt file. You can read more about the robots.txt Robots Exclusion Standard, but we thought we&#8217;d answer some common questions here. What is a robots.txt file? A robots.txt file provides restrictions [&hellip;]<\/p>\n","protected":false},"menu_order":86797,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23783","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23783","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23783\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26415"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23783"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23783"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23783"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}