Arabic Is Not a Low-Resource Language: Structural Load, Register Collapse, and the Misdiagnosis of High-Stakes Text
Arabic is routinely described as a low-resource language. The description is repeated in funding documents, in model cards, and in the framing of research programmes, and it is wrong in a way that shapes what gets built. Arabic has one of the largest continuously produced textual traditions in human history and an enormous contemporary corpus. What it lacks is not text. It is structured annotation of the text it has, and the two are different scarcities requiring different remedies. This article proposes a framework built on that distinction. The Resource Matrix separates volume from annotation depth and places Arabic where it belongs: data-rich and annotation-poor, a quadrant the low-resource label does not describe. Structural Load names the share of meaning a text carries in its structure rather than in its lexis, and the article argues that Arabic's most consequential texts sit at the high end of that scale — which is precisely where general-purpose systems trained on volume perform least reliably. Three further constructs follow. Register Collapse is the treatment of classical, scholarly, standard, and spoken Arabic as a single language in a corpus, when they differ in structure, in function, and in what an error means. Consequence Asymmetry is the observation that competence and stakes run in opposite directions: the texts where an undetected error costs most are the texts a system handles worst. Boundary provenance is the record of where a structural unit begins and ends and on whose authority — cheap to maintain at segmentation, impossible to reconstruct afterwards. The article identifies four predicted failure modes and argues that the costliest is the least visible. Register migration — a question belonging to one register answered fluently in the terms of another — produces output in which nothing is checkably false, and is undetectable by exactly the users most likely to encounter it: those who asked because they did not already know. This is a framework contribution rather than an empirical study, and says so throughout. It cites no external sources. Every proportion in every figure is stipulated to display a shape rather than report a measurement, every classification is argued in the text, and for each construct the article states what observation would refute it. The framework derives from the author's own practice segmenting and coding structurally dense Arabic text, building boundary registries, and producing bilingual scholarly material where register fidelity is a shipping requirement; Section 6 describes that standpoint and its limits openly so the reader can calibrate. The second half of the article is practical. A Structural Adequacy Protocol establishes what kind of deployment a reader has. An eight-test battery — runnable in an afternoon by one competent reader, with no infrastructure — establishes how a system actually behaves on their material, with each test written so that a failure is recognisable rather than requiring expert adjudication. Deployment decision rules map structural load against error cost onto four quadrants and say what each requires. A remediation ladder orders seven interventions by cost, and notes that the first three are close to free while one of them has a deadline: boundary provenance is cheap at the moment of segmentation and unreconstructable afterwards. A final table sets out what to do by role, for builders, buyers, institutions, scholars, funders, and researchers. A worked deployment scenario shows an institution making eight individually defensible decisions and arriving at a configuration where every protocol check returns the adverse value — with no error rate visible, because nobody is checking, and no complaint signal, because the users who would notice a register migration are not the ones asking. A Structural Adequacy Protocol is proposed for anyone deploying a system on Arabic material, together with a boundary provenance record and a worked deployment scenario in which every check returns the adverse value without any decision having been negligent. The practical conclusion is that the remedy for Arabic is not more text and not a larger model. It is annotation, register tagging, and boundary provenance — and the scarce input is the time of scholars who can read the classical registers, a labour that currently falls between every available funding category. Paper 7 of 10 in The Answerability Series. Manuscript ID ARAB-STRUCT-2026-07. 45 pages, 10 figures, 24 tables, five appendices including a protocol worksheet, a boundary provenance record, a construct reference card, and a worked deployment scenario — all released under CC BY 4.0.
Authors
- Syed Shahzad (ORCID: https://orcid.org/0009-0001-7323-1577)
Institutions
- Sir Syed University of Engineering and Technology (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22844717
- Primary Topic
- Authorship Attribution and Profiling
- Type
- preprint