{"id":23827,"date":"2006-09-19T00:00:00","date_gmt":"2006-09-19T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/debugging-blocked-urls\/"},"modified":"2006-09-19T00:00:00","modified_gmt":"2006-09-19T00:00:00","slug":"debugging-blocked-urls","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/debugging-blocked-urls\/","title":{"rendered":"Debugging blocked URLs"},"content":{"rendered":"<p class=\"gargardate\">Tuesday, September 19, 2006<\/p>\n<p>\n  Vanessa&#8217;s been posting a lot lately, and I&#8217;m starting to feel left out. So here my tidbit of<br \/>\n  wisdom for you:  I&#8217;ve noticed a couple of webmasters confused by<br \/>\n  &#8220;<a href=\"https:\/\/support.google.com\/webmasters\/answer\/7440203#submitted_but_blocked\" class=\"external-link\">blocked by robots.txt<\/a>&#8221;<br \/>\n  errors, and I wanted to share the steps I take when debugging<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">robots.txt<\/a> problems:\n<\/p>\n<h2 id=\"a-handy-checklist-for-debugging-a-blocked-url\" tabindex=\"-1\">A handy checklist for debugging a blocked URL<\/h2>\n<p>\n  Let&#8217;s assume you are looking at crawl errors for your website and notice a URL restricted by<br \/>\n  robots.txt that you weren&#8217;t intending to block:\n<\/p>\n<div><\/div>\n<h2 id=\"check-the-robots.txt-analysis-tool\" tabindex=\"-1\">Check the robots.txt analysis tool<\/h2>\n<p>\n  The first thing you should do is go to the<br \/>\n  <a href=\"https:\/\/www.google.com\/webmasters\/tools\/robots-testing-tool\" class=\"external-link\">robots.txt analysis tool<\/a><br \/>\n  for that site. Make sure you are looking at the correct site for that URL, paying attention that<br \/>\n  you are looking at the right protocol and subdomain. (Subdomains and protocols may have their own<br \/>\n  robots.txt file, so <code>https:\/\/www.example.com\/robots.txt<\/code> may be different from<br \/>\n  <code>https:\/\/example.com\/robots.txt<\/code> and may be different from<br \/>\n  <code>https:\/\/amanda.example.com\/robots.txt.<\/code>) Paste the blocked URL into the &#8220;Test URLs<br \/>\n  against this robots.txt file&#8221; box. If the tool reports that it is blocked, you&#8217;ve found your<br \/>\n  problem. If the tool reports that it&#8217;s allowed, we need to investigate further.\n<\/p>\n<p>\n  At the top of the robots.txt analysis tool, take a look at the<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/how-http-status-codes-affect-googles-crawlers\/\">HTTP status code<\/a>.<br \/>\n  If we are reporting anything other than a <code>200 (Success)<\/code> or a<br \/>\n  <code>404 (Not found)<\/code> then we may not be able to reach your robots.txt file, which stops<br \/>\n  our crawling process. (Note that you can see the last time we downloaded your robots.txt file at<br \/>\n  the top of this tool. If you make changes to your file, check this date and time to see if your<br \/>\n  changes were made after our last download.)\n<\/p>\n<h2 id=\"check-for-changes-in-your-robots.txt-file\" tabindex=\"-1\"> Check for changes in your robots.txt file<\/h2>\n<p>\n  If these look fine, you may want to check and see if your robots.txt file has changed since the<br \/>\n  error occurred by checking the date to see when your robots.txt file was last modified. If it<br \/>\n  was modified after the date given for the error in the crawl errors, it might be that someone<br \/>\n  has changed the file so that the new version no longer blocks this URL.\n<\/p>\n<h2 id=\"check-for-redirects-of-the-url\" tabindex=\"-1\"> Check for redirects of the URL<\/h2>\n<p>\n  If you can be certain that this URL isn&#8217;t blocked, check to see if the URL redirects to another<br \/>\n  page. When Googlebot fetches a URL, it checks the robots.txt file to make sure it is allowed to<br \/>\n  access the URL. If the robots.txt file allows access to the URL, but the URL returns a redirect,<br \/>\n  Googlebot checks the robots.txt file again to see if the destination URL is accessible. If at any<br \/>\n  point Googlebot is redirected to a blocked URL, it reports that it could not get the content of<br \/>\n  the original URL because it was blocked by robots.txt.\n<\/p>\n<p>\n  Sometimes this behavior is easy to spot because a particular URL always redirects to another<br \/>\n  one. But sometimes this can be tricky to figure out. For instance:\n<\/p>\n<ul>\n<li>\n   Your site may not have a robots.txt file at all (and therefore, allows access to all pages),<br \/>\n   but a URL on the site may redirect to a different site, which does have a robots.txt file. In<br \/>\n   this case, you may see URLs blocked by robots.txt for your site (even though you don&#8217;t have a<br \/>\n   robots.txt file).\n <\/li>\n<li>\n   Your site may prompt for registration after a certain number of page views. You may have the<br \/>\n   registration page blocked by a robots.txt file. In this case, the URL itself may not redirect,<br \/>\n   but if Googlebot triggers the registration prompt when accessing the URL, it will be redirected<br \/>\n   to the blocked registration page, and the original URL will be listed in the crawl errors page<br \/>\n   as blocked by robots.txt.\n <\/li>\n<\/ul>\n<h2 id=\"ask-for-help\" tabindex=\"-1\">Ask for help<\/h2>\n<p>\n  Finally, if you still can&#8217;t pinpoint the problem, you might want to post on our<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/community\" class=\"external-link\">forum<\/a><br \/>\n  for help. Be sure to include the URL that is blocked in your message. Sometimes its easier for<br \/>\n  other people to notice oversights you may have missed.\n<\/p>\n<p>\n  Good luck debugging! And by the way&mdash;unrelated to robots.txt&mdash;make sure that you don&#8217;t<br \/>\n  have <code>noindex<\/code> <code>meta<\/code> tags at the top of your web pages; those also result in Google not<br \/>\n  showing a web site in our index.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Tuesday, September 19, 2006 Vanessa&#8217;s been posting a lot lately, and I&#8217;m starting to feel left out. So here my tidbit of wisdom for you: I&#8217;ve noticed a couple of webmasters confused by &#8220;blocked by robots.txt&#8221; errors, and I wanted to share the steps I take when debugging robots.txt problems: A handy checklist for debugging [&hellip;]<\/p>\n","protected":false},"menu_order":86590,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23827","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23827","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23827\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26459"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23827"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23827"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23827"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}