Access-Bounded Benchmarks: Evaluating LLMs Where Scaling Cannot Close the Gap (with Khwarezm-100)

We introduce KHWAREZM-100, an evaluation benchmark for large language models (LLMs) built from primary sources of the medieval Khwarezm mathematical school (al-Khwarizmi, al-Biruni, al-Kashi), and use it to argue for a previously under-recognized class of LLM evaluations we term access-bounded benchmarks: tasks on which the frontier gap is bounded below by institutional and linguistic access to training data, not by model scale or compute. Our central contribution is a formal one. We model the frontier laboratory's corpus-acquisition decision as a budget-constrained allocation problem and prove, under a support-conditioned scaling law and a non-substitutability assumption, that the resulting benchmark gap admits a lower bound delta > 0 that is uniform in compute (Theorem 1); no finite compute multiplier closes it (Corollary 1). The proof separates two mechanisms that the informal literature conflates - hard bounds, where the corpus does not exist in machine-readable form, and soft bounds, where it exists but its acquisition is never optimal for the frontier - and this separation resolves the apparent paradox of releasing an access-bounded resource openly (Section 2.4). Three further contributions follow. (i) We instantiate the category with KHWAREZM-100, drawn from the Kitab al-jabr wa-l-muqabala and related corpora, scoring each item jointly on solution correctness and methodological fidelity - adherence to the historical six-case taxonomy and completing-the-square geometry - a dimension absent from standard mathematical benchmarks and made concrete by two worked derivations (Appendix E). (ii) We specify the estimation protocol quantitatively: a closed form for the corpus-sufficiency point alpha*, and a power analysis establishing that the fame-fidelity correlation claim requires n >= 33 source-locked items, which our pilot (n = 15) does not meet (Section 4). (iii) We propose a tri-lingual sentence-aligned (Arabic-Russian-Uzbek) TEI-XML corpus of these sources for targeted pretraining. As preliminary empirical content we report partial construction and a single-model correctness probe with a blind, no-LLM-judge harness; the multi-model, human-scored fidelity study is in progress. Changes in v2 (30 September 2026). The cost hypothesis of Proposition 1 is weakened: j-prime need cost no more than the niche corpus plus the acquisition budget the optimum leaves unspent. The proof already established exactly that inequality, so no result changes. The monotonicity clause is now stated for the value inequality, which is the one that depends on the token budget. Section 2.4 says explicitly that when the corpus is free the cost half of the exchange condition is carried by the unspent budget, and that what keeps cheap domains out of the acquired set is the mixture floor, which admits at most 1/epsilon acquired domains, not the budget constraint.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23060373
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Access-Bounded Benchmarks: Evaluating LLMs Where Scaling Cannot Close the Gap (with Khwarezm-100)

Ravil Akhtyamov
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

Access-Bounded Benchmarks: Evaluating LLMs Where Scaling Cannot Close the Gap (with Khwarezm-100)

Ravil Akhtyamov
preprint en

Abstract

We introduce KHWAREZM-100, an evaluation benchmark for large language models (LLMs) built from primary sources of the medieval Khwarezm mathematical school (al-Khwarizmi, al-Biruni, al-Kashi), and use it to argue for a previously under-recognized class of LLM evaluations we term access-bounded benchmarks: tasks on which the frontier gap is bounded below by institutional and linguistic access to training data, not by model scale or compute. Our central contribution is a formal one. We model the frontier laboratory's corpus-acquisition decision as a budget-constrained allocation problem and prove, under a support-conditioned scaling law and a non-substitutability assumption, that the resulting benchmark gap admits a lower bound delta > 0 that is uniform in compute (Theorem 1); no finite compute multiplier closes it (Corollary 1). The proof separates two mechanisms that the informal literature conflates - hard bounds, where the corpus does not exist in machine-readable form, and soft bounds, where it exists but its acquisition is never optimal for the frontier - and this separation resolves the apparent paradox of releasing an access-bounded resource openly (Section 2.4). Three further contributions follow. (i) We instantiate the category with KHWAREZM-100, drawn from the Kitab al-jabr wa-l-muqabala and related corpora, scoring each item jointly on solution correctness and methodological fidelity - adherence to the historical six-case taxonomy and completing-the-square geometry - a dimension absent from standard mathematical benchmarks and made concrete by two worked derivations (Appendix E). (ii) We specify the estimation protocol quantitatively: a closed form for the corpus-sufficiency point alpha*, and a power analysis establishing that the fame-fidelity correlation claim requires n >= 33 source-locked items, which our pilot (n = 15) does not meet (Section 4). (iii) We propose a tri-lingual sentence-aligned (Arabic-Russian-Uzbek) TEI-XML corpus of these sources for targeted pretraining. As preliminary empirical content we report partial construction and a single-model correctness probe with a blind, no-LLM-judge harness; the multi-model, human-scored fidelity study is in progress. Changes in v2 (30 September 2026). The cost hypothesis of Proposition 1 is weakened: j-prime need cost no more than the niche corpus plus the acquisition budget the optimum leaves unspent. The proof already established exactly that inequality, so no result changes. The monotonicity clause is now stated for the value inequality, which is the one that depends on the token budget. Section 2.4 says explicitly that when the corpus is free the cost half of the exchange condition is carried by the unspent budget, and that what keeps cheap domains out of the acquired set is the mixture floor, which admits at most 1/epsilon acquired domains, not the budget constraint.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.