Describing the Yandex Top Without Ordering It: On-Page Features of the Top 20 and a Technical Profile of the Top 10 Against Positions 31-49

Preprint with data and code: two observational studies of Russian-language Yandex results that test whether the page properties checked by SEO audits differ between pages that rank higher and pages that rank lower. Part I, on-page. 1 538 query–document observations (1 057 unique pages, 97 queries, 562 domains) from the Yandex top 20, region Russia, captured on 26 August 2026; the corpus of Study 2 of doi:10.5281/zenodo.22114786. More than forty page and domain features are summarised as the share of within-query page pairs in which the page with the larger value ranks higher (50 = coin flip). Query words in the title, h1, description and subheadings, query density, title length, heading count, text length, BM25, structured data, images, alt text, lists, links, URL shape, readability and three layers of an AI-text detector fall between 46.0 and 53.9 pairs out of 100. Domain Rating and Tranco popularity reach 55.1 and 55.3. Feature pairs known to be related reach 91.6 and 86.8 on the same measure, so the values near 50 are not an artefact of the measure. Part II, technical. 1 215 clean pages from positions 1–10 against 1 370 clean pages from positions 31, 33, …, 49; 160 queries in 16 niches, Moscow, captured on 9 September 2026, after removing HTTP-200 pages that served bot-protection stubs. The median top-10 HTML document is 157 075 bytes larger, the gap appears in 119 of 160 queries and in all 16 niches, and a cold-cache browser transfers 3 266.7 against 2 159.4 KB. Among the 111 hosts with pages in both groups the gap is 213 875 bytes unpaired and 11 463 bytes when pages are paired within host (sign test p = 0.128): the technical difference separates sites, not pages of one site. A faster server response at the top (−80.5 ms) did not reproduce in the browser measurement (+14.5 ms, interval including zero). The top 10 is also harder to crawl: 17.8% of its URLs did not return HTTP 200 to an external crawler against 8.5% of the control group. Contents. The preprint (PDF and LaTeX source with generated tables and figures); SERP snapshots (query, position, URL, domain); per-observation on-page features; per-page technical features with cleaning flags; host-level robots.txt and protocol probes; browser and stealth-browser measurements; result files; outputs of the original analysis scripts; and code. code/verify_headline_numbers.py recomputes every reported number from the CSV files and exits with code 0 when all values reproduce (standard-library Python). Not included. Page HTML, extracted text, titles and descriptions of the ranked third-party pages. Pages are identified by URL, query and position. Both studies were first published as Russian-language articles on sk-seo.ru; Appendix B of the preprint lists where those articles differ from the data. Text, figures and data: CC BY 4.0. Code: MIT.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22796066
Citations
2
Primary Topic
Misinformation and Its Impacts
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Describing the Yandex Top Without Ordering It: On-Page Features of the Top 20 and a Technical Profile of the Top 10 Against Positions 31-49

Stanislav Kirichenko
2 citations
Zenodo (CERN European Organization for Nuclear Research)
Misinformation and Its Impacts
preprint

Describing the Yandex Top Without Ordering It: On-Page Features of the Top 20 and a Technical Profile of the Top 10 Against Positions 31-49

Stanislav Kirichenko
preprint en
2 citations

Abstract

Preprint with data and code: two observational studies of Russian-language Yandex results that test whether the page properties checked by SEO audits differ between pages that rank higher and pages that rank lower. Part I, on-page. 1 538 query–document observations (1 057 unique pages, 97 queries, 562 domains) from the Yandex top 20, region Russia, captured on 26 August 2026; the corpus of Study 2 of doi:10.5281/zenodo.22114786. More than forty page and domain features are summarised as the share of within-query page pairs in which the page with the larger value ranks higher (50 = coin flip). Query words in the title, h1, description and subheadings, query density, title length, heading count, text length, BM25, structured data, images, alt text, lists, links, URL shape, readability and three layers of an AI-text detector fall between 46.0 and 53.9 pairs out of 100. Domain Rating and Tranco popularity reach 55.1 and 55.3. Feature pairs known to be related reach 91.6 and 86.8 on the same measure, so the values near 50 are not an artefact of the measure. Part II, technical. 1 215 clean pages from positions 1–10 against 1 370 clean pages from positions 31, 33, …, 49; 160 queries in 16 niches, Moscow, captured on 9 September 2026, after removing HTTP-200 pages that served bot-protection stubs. The median top-10 HTML document is 157 075 bytes larger, the gap appears in 119 of 160 queries and in all 16 niches, and a cold-cache browser transfers 3 266.7 against 2 159.4 KB. Among the 111 hosts with pages in both groups the gap is 213 875 bytes unpaired and 11 463 bytes when pages are paired within host (sign test p = 0.128): the technical difference separates sites, not pages of one site. A faster server response at the top (−80.5 ms) did not reproduce in the browser measurement (+14.5 ms, interval including zero). The top 10 is also harder to crawl: 17.8% of its URLs did not return HTTP 200 to an external crawler against 8.5% of the control group. Contents. The preprint (PDF and LaTeX source with generated tables and figures); SERP snapshots (query, position, URL, domain); per-observation on-page features; per-page technical features with cleaning flags; host-level robots.txt and protocol probes; browser and stealth-browser measurements; result files; outputs of the original analysis scripts; and code. code/verify_headline_numbers.py recomputes every reported number from the CSV files and exits with code 0 when all values reproduce (standard-library Python). Not included. Page HTML, extracted text, titles and descriptions of the ranked third-party pages. Pages are identified by URL, query and position. Both studies were first published as Russian-language articles on sk-seo.ru; Appendix B of the preprint lists where those articles differ from the data. Text, figures and data: CC BY 4.0. Code: MIT.

Zenodo (CERN European Organization for Nuclear Research)
Misinformation and Its Impacts
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.