Describing the Yandex Top Without Ordering It: On-Page Features of the Top 20 and a Technical Profile of the Top 10 Against Positions 31-49
Preprint with data and code: two observational studies of Russian-language Yandex results that test whether the page properties checked by SEO audits differ between pages that rank higher and pages that rank lower. Part I, on-page. 1 538 query–document observations (1 057 unique pages, 97 queries, 562 domains) from the Yandex top 20, region Russia, captured on 26 August 2026; the corpus of Study 2 of doi:10.5281/zenodo.22114786. More than forty page and domain features are summarised as the share of within-query page pairs in which the page with the larger value ranks higher (50 = coin flip). Query words in the title, h1, description and subheadings, query density, title length, heading count, text length, BM25, structured data, images, alt text, lists, links, URL shape, readability and three layers of an AI-text detector fall between 46.0 and 53.9 pairs out of 100. Domain Rating and Tranco popularity reach 55.1 and 55.3. Feature pairs known to be related reach 91.6 and 86.8 on the same measure, so the values near 50 are not an artefact of the measure. Part II, technical. 1 215 clean pages from positions 1–10 against 1 370 clean pages from positions 31, 33, …, 49; 160 queries in 16 niches, Moscow, captured on 9 September 2026, after removing HTTP-200 pages that served bot-protection stubs. The median top-10 HTML document is 157 075 bytes larger, the gap appears in 119 of 160 queries and in all 16 niches, and a cold-cache browser transfers 3 266.7 against 2 159.4 KB. Among the 111 hosts with pages in both groups the gap is 213 875 bytes unpaired and 11 463 bytes when pages are paired within host (sign test p = 0.128): the technical difference separates sites, not pages of one site. A faster server response at the top (−80.5 ms) did not reproduce in the browser measurement (+14.5 ms, interval including zero). The top 10 is also harder to crawl: 17.8% of its URLs did not return HTTP 200 to an external crawler against 8.5% of the control group. Contents. The preprint (PDF and LaTeX source with generated tables and figures); SERP snapshots (query, position, URL, domain); per-observation on-page features; per-page technical features with cleaning flags; host-level robots.txt and protocol probes; browser and stealth-browser measurements; result files; outputs of the original analysis scripts; and code. code/verify_headline_numbers.py recomputes every reported number from the CSV files and exits with code 0 when all values reproduce (standard-library Python). Not included. Page HTML, extracted text, titles and descriptions of the ranked third-party pages. Pages are identified by URL, query and position. Both studies were first published as Russian-language articles on sk-seo.ru; Appendix B of the preprint lists where those articles differ from the data. Text, figures and data: CC BY 4.0. Code: MIT.
Authors
- Stanislav Kirichenko (ORCID: https://orcid.org/0009-0001-2914-9541)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-16
- DOI
- https://doi.org/10.5281/zenodo.22796066
- Citations
- 2
- Primary Topic
- Misinformation and Its Impacts
- Type
- preprint