Access-Bounded Benchmarks: Evaluating LLMs Where Scaling Cannot Close the Gap (with Khwarezm-100)
We introduce KHWAREZM-100, an evaluation benchmark for large language models (LLMs) built from primary sources of the medieval Khwarezm mathematical school (al-Khwarizmi, al-Biruni, al-Kashi), and use it to argue for a previously under-recognized class of LLM evaluations we term access-bounded benchmarks: tasks on which the frontier gap is bounded below by institutional and linguistic access to training data, not by model scale or compute. Our central contribution is a formal one. We model the frontier laboratory's corpus-acquisition decision as a budget-constrained allocation problem and prove, under a support-conditioned scaling law and a non-substitutability assumption, that the resulting benchmark gap admits a lower bound delta > 0 that is uniform in compute (Theorem 1); no finite compute multiplier closes it (Corollary 1). The proof separates two mechanisms that the informal literature conflates - hard bounds, where the corpus does not exist in machine-readable form, and soft bounds, where it exists but its acquisition is never optimal for the frontier - and this separation resolves the apparent paradox of releasing an access-bounded resource openly (Section 2.4). Three further contributions follow. (i) We instantiate the category with KHWAREZM-100, drawn from the Kitab al-jabr wa-l-muqabala and related corpora, scoring each item jointly on solution correctness and methodological fidelity - adherence to the historical six-case taxonomy and completing-the-square geometry - a dimension absent from standard mathematical benchmarks and made concrete by two worked derivations (Appendix E). (ii) We specify the estimation protocol quantitatively: a closed form for the corpus-sufficiency point alpha*, and a power analysis establishing that the fame-fidelity correlation claim requires n >= 33 source-locked items, which our pilot (n = 15) does not meet (Section 4). (iii) We propose a tri-lingual sentence-aligned (Arabic-Russian-Uzbek) TEI-XML corpus of these sources for targeted pretraining. As preliminary empirical content we report partial construction and a single-model correctness probe with a blind, no-LLM-judge harness; the multi-model, human-scored fidelity study is in progress. Changes in v2 (30 September 2026). The cost hypothesis of Proposition 1 is weakened: j-prime need cost no more than the niche corpus plus the acquisition budget the optimum leaves unspent. The proof already established exactly that inequality, so no result changes. The monotonicity clause is now stated for the value inequality, which is the one that depends on the token budget. Section 2.4 says explicitly that when the corpus is free the cost half of the exchange condition is carried by the unspent budget, and that what keeps cheap domains out of the acquired set is the mixture floor, which admits at most 1/epsilon acquired domains, not the budget constraint.
Authors
- Ravil Akhtyamov
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23060373
- Primary Topic
- Topic Modeling
- Type
- preprint