Yandex describes its own pipeline in four stages, and knowing which stage a problem belongs to is what turns a vague complaint into a fixable one.
Stage 1 — crawling
The robot decides which sites to crawl, how often, and how many pages. Its list of known pages comes from internal and external links, the sitemap, Yandex Metrica data, and robots.txt. Pages larger than 10 MB are not indexed.
Response handling is explicit:
200 OK— crawled.3XX— the robot crawls the redirect target.4XXand5XX— the page is not included, and is removed if it was already in search.
And the useful trick that follows: if a page is temporarily broken through a CMS fault, configure the server to answer 429 instead. The robot checks the code and does not drop the page. But sustained 429 tells Yandex the server is struggling and reduces the crawl rate, so it is a bandage, not a setting.
HTTP/2 is supported — faster loading and lower server load, though Yandex is clear it does not change crawl frequency or ranking, and HTTP/1.1 continues to work.
Stage 2 — indexing
The robot analyses and stores the content: the description meta tag, the title, Schema.org markup for the snippet; the noindex directive, which keeps the page out; rel="canonical", declaring the preferred address for a group; and the text, images and video. Where content matches across pages, they become duplicates.
Stages 3 and 4
A database of pages eligible for search is assembled, and results are generated from it. Results update regularly, and those updates can move a site’s position without anything on the site changing — which is the correct answer to most “we dropped and nothing changed” questions.
Why the model matters
Each stage has its own diagnostic. Not crawled is a links, sitemap or availability problem. Crawled but not indexed is a directive, duplicate or quality problem. Indexed but not ranking is a relevance problem. Working the wrong stage is the most expensive mistake in technical SEO, and this page is the map that prevents it.
Comment nous l'appliquons
The 429 behaviour is the practical gift here: during a CMS incident it keeps pages in the index instead of letting them be dropped for 5XX responses, which turns a two-hour outage into a non-event rather than a two-week recovery. We put it in the incident runbook for every client with a fragile stack — and we take it out again once the incident ends, because sustained 429 costs crawl rate.
Services associés