Benchmark-Driven Selection of AI Can Improve Capabilities of Reasoning Language Models in Epistemology
Abstract Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before task completion and that the traces can sometimes help them to generalize better than past models. We show that better performance results not only from test time algorithmic improvements or model sizes but also from letting tasks from impactful benchmarks inspire curricula for learning. We call this benchmark-driven selection of AI and show its effects on DeepSeek-R1, the first open-weight large reasoning model, using a novel sequential decision-making problem that we contributed to the philosophy category of Humanity’s Last Exam, a frontier academic AI benchmark. Steering development of AI by impactful benchmarks trades evaluation for learning and makes novelty of test tasks key for measuring generalization capabilities of reasoning models. Consequently, some benchmarks inspire curricula for post-training that improve external validity of the models. Understanding public benchmarks as mere test sets risks confusing unseen tasks measuring uncontaminated generalization with known tasks that improve external validity. This paper explores new avenues for experimentation in epistemology and shows that formal AI evaluation problems are helpful for understanding model capabilities and their evolution.
Authors
- Petr Špelda (ORCID: https://orcid.org/0000-0003-4199-645X)
- Vít Střítecký (ORCID: https://orcid.org/0000-0003-1778-3657)
Institutions
- Charles University (CZ)
Publication Details
- Journal
- Episteme
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1017/epi.2026.10139
- Primary Topic
- Logic, Reasoning, and Knowledge
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Univerzita Karlova v Praze