A sitemap tells search engines the site’s current structure. Yandex supports XML (recommended, because it carries extra information) and TXT (a plain list of URLs). It does not guarantee that the URLs will appear in results.
When you actually need one
Yandex is candid that its algorithms usually find pages through internal and external links, but the robot can miss some. A sitemap is worthwhile when the site has many pages, pages with no navigation links, or a deeply nested structure. A small, well-linked site gains little from one.
The XML elements
loc— required, the page address.lastmod— last update date.changefreq— how often the page changes.priority— 0.0 to 1.0, and Yandex explains what it does: the robot loads pages sorted by the presence and value of this coefficient. It orders crawling. Set it for the URLs that genuinely matter, and it means nothing if every page is 1.0.
Each optional element has a maximum size of 100 bytes.
Requirements
- UTF-8; Cyrillic URLs accepted in original or encoded form.
- 50,000 links maximum — split across files with an index file beyond that.
- 50 MB uncompressed.
- Only URLs from the domain the file sits on, and the file must be on that domain.
- The file must return
200 OK.
RSS and Atom feeds are not supported as sitemaps.
The step people skip
Yandex’s procedure begins with defining canonical URLs for the pages that will be listed. A sitemap full of duplicate variants tells the robot to crawl work you have already decided is redundant. Then create the file, validate it, and declare it in robots.txt or in the console.
For large sites
Drop pages Yandex already knows and keep the file for new and frequently updated ones, marking volatile files with lastmod in the index file. Crawl statistics shows what is already known — which turns the sitemap from an inventory into a change feed.
How we apply this
The priority element is the part most implementations waste: it orders the crawl queue, so a generator emitting 0.8 for everything has thrown away a real lever. We set it deliberately on the pages that matter and leave the rest unset. And we start with canonicals, because a sitemap listing every parameter variant is asking the robot to spend its budget on pages we have already ruled out.
Related services