{"id":28221,"date":"2026-09-08T21:52:59","date_gmt":"2026-09-08T18:52:59","guid":{"rendered":"https:\/\/alienroad.com\/baidu-bilgi-bankasi\/robots-txt-and-baiduspider\/"},"modified":"2026-09-09T02:55:31","modified_gmt":"2026-09-08T23:55:31","slug":"robots-txt-and-baiduspider","status":"publish","type":"ar_bkb","link":"https:\/\/alienroad.com\/baidu-bilgi-bankasi\/robots-txt-and-baiduspider\/","title":{"rendered":"robots.txt and Baiduspider"},"content":{"rendered":"<p>Baiduspider checks <code>robots.txt<\/code> in the root of the domain before crawling. The syntax follows the standard protocol; the details below are where Baidu differs or is more explicit than most engines.<\/p>\n<h2>The rule most sites get wrong<\/h2>\n<p>Baidu states it plainly: <strong>if you want everything indexed, do not create a robots.txt file at all.<\/strong> The file exists only to exclude. An empty or absent file leaves the site fully open, which is the correct configuration for most sites and the one people replace with a half-remembered template.<\/p>\n<h2>Blocking is not removal<\/h2>\n<p>A page blocked in <code>robots.txt<\/code> <strong>can still appear in Baidu results<\/strong> when other sites link to it. Its content is not crawled, indexed or displayed \u2014 what appears is other sites&#8217; description of it. Blocking controls crawling; it does not control presence.<\/p>\n<h2>Directive precision<\/h2>\n<ul>\n<li><code>Disallow: \/help<\/code> blocks <code>\/help.html<\/code>, <code>\/helpabc.html<\/code> <em>and<\/em> <code>\/help\/index.html<\/code>.<\/li>\n<li><code>Disallow: \/help\/<\/code> blocks only <code>\/help\/index.html<\/code> \u2014 the first two remain crawlable.<\/li>\n<li>An empty <code>Disallow:<\/code> permits everything; at least one <code>Disallow<\/code> record is required in a valid file.<\/li>\n<li><strong>Order matters:<\/strong> the robot applies <em>the first matching <code>Allow<\/code> v\u0259 ya <code>Disallow<\/code> line<\/em>. This is not the longest-match rule other engines use, and a file written for Google can behave differently here.<\/li>\n<li>Wildcards: <code>*<\/code> matches any sequence, <code>$<\/code> matches end of line.<\/li>\n<li><strong>Matching is case-sensitive and exact.<\/strong> Baidu warns that a case mismatch makes the rule ineffective.<\/li>\n<\/ul>\n<h2>Baidu-specific meta directives<\/h2>\n<ul>\n<li><code>&lt;meta name=\"robots\" content=\"nofollow\"&gt;<\/code> \u2014 do not follow links or pass value; <code>&lt;meta name=\"Baiduspider\" content=\"nofollow\"&gt;<\/code> restricts this to Baidu alone.<\/li>\n<li><code>&lt;meta name=\"robots\" content=\"noarchive\"&gt;<\/code> \u2014 no cached snapshot; the Baidu-only form is <code>&lt;meta name=\"Baiduspider\" content=\"noarchive\"&gt;<\/code>. Note that <strong>noarchive suppresses the snapshot only<\/strong> \u2014 the page is still indexed and still shown with a summary.<\/li>\n<li>Per-link control with <code>rel=\"nofollow\"<\/code> is supported.<\/li>\n<\/ul>\n<h2>Identifying the real Baiduspider<\/h2>\n<p>Baidu does not publish its IP ranges; they change. Verification is by user agent plus reverse DNS, and there are three user-agent families \u2014 <strong>mobile<\/strong>, <strong>PC<\/strong>, and <strong>mini-program<\/strong>, with <code>Baiduspider-render<\/code> variants for rendering. Blocking by IP is not a supported strategy; checking the agent first and confirming by DNS is.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Baiduspider checks robots.txt in the root of the domain before crawling. The syntax follows the standard protocol; the details below are where Baidu differs or is more explicit than most engines. The rule most sites get wrong Baidu states it plainly: if you want everything indexed, do not create a robots.txt file at all. The [&hellip;]<\/p>\n","protected":false},"menu_order":9,"template":"","meta":{"footnotes":""},"ar_bkb_kategori":[744],"ar_bkb_etiket":[],"class_list":["post-28221","ar_bkb","type-ar_bkb","status-publish","has-post-thumbnail","hentry","ar_bkb_kategori-indexing-and-submission"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_bkb\/28221","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_bkb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_bkb"}],"version-history":[{"count":2,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_bkb\/28221\/revisions"}],"predecessor-version":[{"id":29220,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_bkb\/28221\/revisions\/29220"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/28985"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=28221"}],"wp:term":[{"taxonomy":"ar_bkb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_bkb_kategori?post=28221"},{"taxonomy":"ar_bkb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_bkb_etiket?post=28221"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}