Site indexing

What files can be included in search results?

Ask AI about this page

2 min read

Yandex indexes HTML plus a defined list of file formats, each up to 10 MB: PDF; Microsoft Office DOC, DOCX, XLS, XLSX, PPT, PPTX; OpenOffice ODT, ODS, ODP, ODG; text formats RTF and TXT; and Flash SWF. Long URLs — too many CGI parameters or nested directories — can prevent indexing regardless of format.

PDF: the rules that decide whether a document exists

  • Only text content is indexed. Text represented as images is not.
  • A PDF containing only images has just its first three pages indexed.
  • A PDF containing text is indexed in full.

That middle rule is worth pausing on. A forty-page scanned catalogue contributes three pages of essentially nothing. For any organisation whose substance sits in PDFs — manufacturers, professional services, public bodies — the difference between a scanned document and a text-based one is the difference between an asset and dead weight. Running OCR over an archive is often the single highest-value indexing job available.

SWF

Indexed when directly linked or embedded with object или embed. Yandex indexes text from DefineText, DefineText2, DefineEditText and Metadata, and links from DoAction, DefineButton and DefineButton2. Useful content inside an SWF can lead the robot back to the page hosting it.

Frames

frameset and frame are permitted: the robot indexes the loaded content and identifies the source document from the frame contents. Note the contrast with iframe, whose loaded documents are not indexed.

What to do with this

Audit which documents carry real value, confirm they are text rather than scans, keep them under 10 MB, and give them clean short URLs. Then treat them as pages: a well-named, text-bearing PDF linked from a relevant page is indexable content, and a 60 MB scan behind a parameterised download URL is not.

Alien Road

Как мы это применяем

The three-page rule for image-only PDFs is the fact we most often use to justify an OCR pass. Clients with technical documentation archives usually assume the whole library is searchable; in practice each scanned document contributes almost nothing. Checking is trivial — try selecting text in the PDF — and the fix converts a dead archive into indexable content.

Связанные услуги

Поделиться

© Copyright 2026 Alien Road. All rights reserved.