Sovereign AI and Epistemic Dependence: Infrastructure, Language, and Who Answers for a Region's Knowledge Layer
Sovereign AI is measured in megawatts, chips, and capital. A region announces a gigawatt campus, secures an export licence for tens of thousands of accelerators, and is described as having achieved compute sovereignty. This article argues that the accounting is measuring the visible layers rather than the determinative ones, and that a polity can own the entire physical stack while remaining dependent for the thing sovereignty was supposed to secure: authoritative answers about its own language, law, history, and practice. The article names that condition Epistemic Dependence and develops a seven-layer Sovereignty Stack — energy, silicon design, fabrication, data centre, foundation model, evaluation, application — arguing that sovereignty is a property of layers rather than of nations, and that no polity is sovereign at every layer or needs to be. Two constructs follow. Layer Mismatch is the concentration of capital in layers that are highly visible and weakly determinative. Conditional Capacity is installed compute that operates under another jurisdiction's licence and is revocable by a decision taken elsewhere. The sixth layer is the central contribution. Evaluation — the capacity to establish whether a system performs adequately on a region's own language, sources, and norms — appears in no sovereignty definition or index reviewed here, attracts a negligible share of capital, and is the layer on which every other layer's value depends. A region that cannot evaluate cannot know whether the models it has bought, built, or licensed serve it. It is also the only layer that is cheap, quick to build, and not revocable by anyone else. Arabic-language capability is used as the instrumented case, and the finding is uncomfortable. A published Arabic benchmarking platform reports that large closed frontier models developed outside the region outperform regional Arabic-centric models by sizable margins, and that model size alone does not predict Arabic performance — tokenisation, corpus size, and Arabic-specific tuning do. At the same time, most regional models in the surveyed field are adaptations of foreign base models rather than systems trained from scratch, which carries dependence forward into the layer meant to resolve it: an adapted model cannot be rebuilt if its base becomes unavailable, and inherits the base's tokenisation. Comparative material covers five national programmes operating at five different layers. An Epistemic Dependence Audit is proposed, together with a Conditional Capacity Ratio that any government could compute from records it already holds and none publishes, and an indicative costing of a national evaluation capacity — a small permanent institution, a rounding error against a single data centre. The article is descriptive rather than prescriptive about geopolitics. It takes no position on any state's export policy, treats conditional capacity as a structural feature applicable symmetrically to every jurisdiction in the current configuration, and states throughout what its evidence can and cannot support. All capacity and capital figures are announced rather than delivered, and the register grades each. Paper 6 of 10 in The Answerability Series. Manuscript ID SOV-EPIST-2026-06. 43 pages, 10 figures, 22 tables, five appendices including an audit worksheet, a ratio computation sheet, a layer reference card, and a register of Arabic models and evaluation resources — all released under CC BY 4.0.
Authors
- Syed Shahzad (ORCID: https://orcid.org/0009-0001-7323-1577)
Institutions
- Sir Syed University of Engineering and Technology (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-18
- DOI
- https://doi.org/10.5281/zenodo.22832211
- Primary Topic
- Economic and Technological Innovation
- Type
- preprint