Baiduspider checks robots.txt in the root of the domain before crawling. The syntax follows the standard protocol; the details below are where Baidu differs or is more explicit than most engines.
The rule most sites get wrong
Baidu states it plainly: if you want everything indexed, do not create a robots.txt file at all. The file exists only to exclude. An empty or absent file leaves the site fully open, which is the correct configuration for most sites and the one people replace with a half-remembered template.
Blocking is not removal
A page blocked in robots.txt can still appear in Baidu results when other sites link to it. Its content is not crawled, indexed or displayed — what appears is other sites’ description of it. Blocking controls crawling; it does not control presence.
Directive precision
Disallow: /helpblocks/help.html,/helpabc.htmland/help/index.html.Disallow: /help/blocks only/help/index.html— the first two remain crawlable.- An empty
Disallow:permits everything; at least oneDisallowrecord is required in a valid file. - Order matters: the robot applies the first matching
AllowoDisallowline. This is not the longest-match rule other engines use, and a file written for Google can behave differently here. - Wildcards:
*matches any sequence,$matches end of line. - Matching is case-sensitive and exact. Baidu warns that a case mismatch makes the rule ineffective.
Baidu-specific meta directives
<meta name="robots" content="nofollow">— do not follow links or pass value;<meta name="Baiduspider" content="nofollow">restricts this to Baidu alone.<meta name="robots" content="noarchive">— no cached snapshot; the Baidu-only form is<meta name="Baiduspider" content="noarchive">. Note that noarchive suppresses the snapshot only — the page is still indexed and still shown with a summary.- Per-link control with
rel="nofollow"is supported.
Identifying the real Baiduspider
Baidu does not publish its IP ranges; they change. Verification is by user agent plus reverse DNS, and there are three user-agent families — mobile, PC, and mini-program, with Baiduspider-render variants for rendering. Blocking by IP is not a supported strategy; checking the agent first and confirming by DNS is.
Cómo lo aplicamos
The first-match rule is the one that bites teams reusing a robots.txt written for Google, where longest-match wins — the same file can permit in one engine and block in the other. And we check case-sensitivity explicitly, because Baidu says outright that a mismatched case makes the rule do nothing, silently.
Servicios relacionados