Reading steppe names: what a transcription channel can and cannot recover from the Chinese record of the Xiongnu

Chinese histories record hundreds of Inner Asian names in characters chosen for their sound, and two thousand years of change in Chinese pronunciation have left the modern Mandarin readings unrelated to what the scribes heard. We model the transcription itself as a noisy channel, train it where both sides are known, and measure what can and cannot be recovered. We build and release three transcription corpora, two of them from sources not previously available in machine-readable form: 9,329 Chinese–Mongolian pairs from the Secret History of the Mongols, 594 Chinese–Uyghur pairs from Ligeti's edition of a Ming glossary, and 1,151 verified Chinese–Sanskrit pairs extracted from a Buddhist dictionary that mixes phonetic spellings with translations; and three medieval Turkic lexicons in machine-readable form. On held-out data, the channel reconstructs 31% of Turkic words exactly and 81% within two segments. We have three main results. One reading of a Xiongnu name survives every test we can apply: 若鞮, the element closing the formal names of six later chanyu, reads as İnaktı "the trusted ones". We clarify a century-old disagreement: the dispute over 冒頓 turns entirely on which value is taken for one character, and we price each branch. We also report a set of measured negative results: language attribution is not supportable on available reference data, since at 閼氏, the chanyu's consort title, one spelling yields a Yeniseian elte "wife" or a Turkic elti "lady" by identical steps; nor is full reconstruction supportable at the density available in the record. We also answer a question that has remained unresolved since 1974. Doerfer argued that chance resemblance is pervasive in long-range comparison; Altmann, reviewing him, sharpened it into a threshold problem: is it a sound law if 49% of cases show one correspondence and 51% another? It cannot be settled by a percentage: the threshold is the baseline of the component being measured, and it moves with how many outcomes that component has. This file contains the paper (pages 1–25) followed by Supplementary Material S1, "Candidate readings of the Xiongnu record" (pages S1-1 to S1-53).

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23178845
Primary Topic
Linguistics and Cultural Studies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Reading steppe names: what a transcription channel can and cannot recover from the Chinese record of the Xiongnu

Aziz Müfit Uluğ
Zenodo (CERN European Organization for Nuclear Research)
Linguistics and Cultural Studies
preprint

Reading steppe names: what a transcription channel can and cannot recover from the Chinese record of the Xiongnu

Aziz Müfit Uluğ
preprint en

Abstract

Chinese histories record hundreds of Inner Asian names in characters chosen for their sound, and two thousand years of change in Chinese pronunciation have left the modern Mandarin readings unrelated to what the scribes heard. We model the transcription itself as a noisy channel, train it where both sides are known, and measure what can and cannot be recovered. We build and release three transcription corpora, two of them from sources not previously available in machine-readable form: 9,329 Chinese–Mongolian pairs from the Secret History of the Mongols, 594 Chinese–Uyghur pairs from Ligeti's edition of a Ming glossary, and 1,151 verified Chinese–Sanskrit pairs extracted from a Buddhist dictionary that mixes phonetic spellings with translations; and three medieval Turkic lexicons in machine-readable form. On held-out data, the channel reconstructs 31% of Turkic words exactly and 81% within two segments. We have three main results. One reading of a Xiongnu name survives every test we can apply: 若鞮, the element closing the formal names of six later chanyu, reads as İnaktı "the trusted ones". We clarify a century-old disagreement: the dispute over 冒頓 turns entirely on which value is taken for one character, and we price each branch. We also report a set of measured negative results: language attribution is not supportable on available reference data, since at 閼氏, the chanyu's consort title, one spelling yields a Yeniseian elte "wife" or a Turkic elti "lady" by identical steps; nor is full reconstruction supportable at the density available in the record. We also answer a question that has remained unresolved since 1974. Doerfer argued that chance resemblance is pervasive in long-range comparison; Altmann, reviewing him, sharpened it into a threshold problem: is it a sound law if 49% of cases show one correspondence and 51% another? It cannot be settled by a percentage: the threshold is the baseline of the component being measured, and it moves with how many outcomes that component has. This file contains the paper (pages 1–25) followed by Supplementary Material S1, "Candidate readings of the Xiongnu record" (pages S1-1 to S1-53).

Zenodo (CERN European Organization for Nuclear Research)
Linguistics and Cultural Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reading steppe names: what a transcription channel can and cannot recover from the Chinese record of the Xiongnu — Aziz Müfit Uluğ · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS