Kaiying Reachable Capability Mapping (KRCM): A Framework for Measuring Conditional Capability Structure in Artificial Intelligence

Artificial-intelligence capability is commonly summarized by endpoint performance on fixed benchmark items. Such scores are useful for ranking systems but are generally insufficient for identifying what transformation a system can perform, under which execution conditions it remains reachable, where a multi-step execution failed, or whether an observed contrast reflects model capability, scaffolding, context, tools, resource allocation, or a particular measurement design. Recent work has begun to move beyond aggregate scores through psychometric modeling, fine-grained ability taxonomies, cognitive diagnosis, adaptive testing, process evaluation, and mechanistic intervention. These developments establish important components of diagnostic evaluation, but they do not by themselves define an integrated measurement object linking task structure, execution process, operating condition, intervention, uncertainty, and the strength of the resulting capability claim. This paper presents Kaiying Reachable Capability Mapping (KRCM), a frozen operational framework for reconstructing conditional capability structure in artificial intelligence. KRCM defines item-level reach as a function of the exact model or system, task instance, registered execution condition, initial or retained process state, and registered stochasticity. Its canonical task-family object is conditional reachability, 𝜌𝑀,𝒟 (Ω), rather than a context-free scalar ability score. The framework separates Representation Sufficiency, local or transformational capability, reachable capability, and observed performance; it likewise separates an observed reasoning trace from reasoning capability and from the final product. Composite tasks may be represented by an item–capability execution graph (ICEG) with explicit admissible paths, node-level provenance, intervention status, and censoring. Candidate capability ontologies and Q-matrices are versioned and provisional: dimensions are proposed by ontology but licensed only by empirical dissociation. Stronger claims require correspondingly stronger evidence, including matched retained-state interventions for causal state dependence, fresh-context local controls that can support only conditional primitive-versus-composition dissociations under registered control conditions, resource matching for mechanism attribution, and identifiability diagnostics for latent parameters. KRCM supports both direct empirical reachability and, when the crossed design permits, restricted multidimensional latent models with condition effects and identified interactions. It also supports adaptive interrogation in which different systems may receive different probes; comparability is supplied by a common measurement geometry, calibrated inference rules, uncertainty standards, and held-out confirmation rather than identical item exposure. A frozen reference discipline treats the model, task bank, execution condition, scoring system, and relevant runtime configuration jointly as a measurement instrument. KRCM-Minimal and KRCM-Full conformance levels distinguish structural compliance from scientific validity. A released empirical companion study provides a worked anchor: the same deterministic local transformation became easier or harder under different multi-step execution regimes across model–task families, while earlier preregistered estimands were retained in the study history after adversarial review showed that they did not identify their intended mechanisms. The empirical result supports conditional reachability, not a universal composition deficit or a specific working-memory mechanism. KRCM therefore reframes AI capability measurement as reconstruction of a versioned, uncertainty-bounded reachable-capability region and frontier, with explicit limits on what the available evidence can identify. Conceptually, the framework separates condition-dependent capability realization from the registered observation and scoring channel: it reconstructs an effective reachable-capability structure over declared probes and conditions, not a unique internal mechanistic state.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22739988
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Kaiying Reachable Capability Mapping (KRCM): A Framework for Measuring Conditional Capability Structure in Artificial Intelligence

Kai Wang
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Kaiying Reachable Capability Mapping (KRCM): A Framework for Measuring Conditional Capability Structure in Artificial Intelligence

Kai Wang
preprint en

Abstract

Artificial-intelligence capability is commonly summarized by endpoint performance on fixed benchmark items. Such scores are useful for ranking systems but are generally insufficient for identifying what transformation a system can perform, under which execution conditions it remains reachable, where a multi-step execution failed, or whether an observed contrast reflects model capability, scaffolding, context, tools, resource allocation, or a particular measurement design. Recent work has begun to move beyond aggregate scores through psychometric modeling, fine-grained ability taxonomies, cognitive diagnosis, adaptive testing, process evaluation, and mechanistic intervention. These developments establish important components of diagnostic evaluation, but they do not by themselves define an integrated measurement object linking task structure, execution process, operating condition, intervention, uncertainty, and the strength of the resulting capability claim. This paper presents Kaiying Reachable Capability Mapping (KRCM), a frozen operational framework for reconstructing conditional capability structure in artificial intelligence. KRCM defines item-level reach as a function of the exact model or system, task instance, registered execution condition, initial or retained process state, and registered stochasticity. Its canonical task-family object is conditional reachability, 𝜌𝑀,𝒟 (Ω), rather than a context-free scalar ability score. The framework separates Representation Sufficiency, local or transformational capability, reachable capability, and observed performance; it likewise separates an observed reasoning trace from reasoning capability and from the final product. Composite tasks may be represented by an item–capability execution graph (ICEG) with explicit admissible paths, node-level provenance, intervention status, and censoring. Candidate capability ontologies and Q-matrices are versioned and provisional: dimensions are proposed by ontology but licensed only by empirical dissociation. Stronger claims require correspondingly stronger evidence, including matched retained-state interventions for causal state dependence, fresh-context local controls that can support only conditional primitive-versus-composition dissociations under registered control conditions, resource matching for mechanism attribution, and identifiability diagnostics for latent parameters. KRCM supports both direct empirical reachability and, when the crossed design permits, restricted multidimensional latent models with condition effects and identified interactions. It also supports adaptive interrogation in which different systems may receive different probes; comparability is supplied by a common measurement geometry, calibrated inference rules, uncertainty standards, and held-out confirmation rather than identical item exposure. A frozen reference discipline treats the model, task bank, execution condition, scoring system, and relevant runtime configuration jointly as a measurement instrument. KRCM-Minimal and KRCM-Full conformance levels distinguish structural compliance from scientific validity. A released empirical companion study provides a worked anchor: the same deterministic local transformation became easier or harder under different multi-step execution regimes across model–task families, while earlier preregistered estimands were retained in the study history after adversarial review showed that they did not identify their intended mechanisms. The empirical result supports conditional reachability, not a universal composition deficit or a specific working-memory mechanism. KRCM therefore reframes AI capability measurement as reconstruction of a versioned, uncertainty-bounded reachable-capability region and frontier, with explicit limits on what the available evidence can identify. Conceptually, the framework separates condition-dependent capability realization from the registered observation and scoring channel: it reconstructs an effective reachable-capability structure over declared probes and conditions, not a unique internal mechanistic state.

Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.