Kaiying Reachable Capability Mapping (KRCM): A Framework for Measuring Conditional Capability Structure in Artificial Intelligence
Artificial-intelligence capability is commonly summarized by endpoint performance on fixed benchmark items. Such scores are useful for ranking systems but are generally insufficient for identifying what transformation a system can perform, under which execution conditions it remains reachable, where a multi-step execution failed, or whether an observed contrast reflects model capability, scaffolding, context, tools, resource allocation, or a particular measurement design. Recent work has begun to move beyond aggregate scores through psychometric modeling, fine-grained ability taxonomies, cognitive diagnosis, adaptive testing, process evaluation, and mechanistic intervention. These developments establish important components of diagnostic evaluation, but they do not by themselves define an integrated measurement object linking task structure, execution process, operating condition, intervention, uncertainty, and the strength of the resulting capability claim. This paper presents Kaiying Reachable Capability Mapping (KRCM), a frozen operational framework for reconstructing conditional capability structure in artificial intelligence. KRCM defines item-level reach as a function of the exact model or system, task instance, registered execution condition, initial or retained process state, and registered stochasticity. Its canonical task-family object is conditional reachability, 𝜌𝑀,𝒟 (Ω), rather than a context-free scalar ability score. The framework separates Representation Sufficiency, local or transformational capability, reachable capability, and observed performance; it likewise separates an observed reasoning trace from reasoning capability and from the final product. Composite tasks may be represented by an item–capability execution graph (ICEG) with explicit admissible paths, node-level provenance, intervention status, and censoring. Candidate capability ontologies and Q-matrices are versioned and provisional: dimensions are proposed by ontology but licensed only by empirical dissociation. Stronger claims require correspondingly stronger evidence, including matched retained-state interventions for causal state dependence, fresh-context local controls that can support only conditional primitive-versus-composition dissociations under registered control conditions, resource matching for mechanism attribution, and identifiability diagnostics for latent parameters. KRCM supports both direct empirical reachability and, when the crossed design permits, restricted multidimensional latent models with condition effects and identified interactions. It also supports adaptive interrogation in which different systems may receive different probes; comparability is supplied by a common measurement geometry, calibrated inference rules, uncertainty standards, and held-out confirmation rather than identical item exposure. A frozen reference discipline treats the model, task bank, execution condition, scoring system, and relevant runtime configuration jointly as a measurement instrument. KRCM-Minimal and KRCM-Full conformance levels distinguish structural compliance from scientific validity. A released empirical companion study provides a worked anchor: the same deterministic local transformation became easier or harder under different multi-step execution regimes across model–task families, while earlier preregistered estimands were retained in the study history after adversarial review showed that they did not identify their intended mechanisms. The empirical result supports conditional reachability, not a universal composition deficit or a specific working-memory mechanism. KRCM therefore reframes AI capability measurement as reconstruction of a versioned, uncertainty-bounded reachable-capability region and frontier, with explicit limits on what the available evidence can identify. Conceptually, the framework separates condition-dependent capability realization from the registered observation and scoring channel: it reconstructs an effective reachable-capability structure over declared probes and conditions, not a unique internal mechanistic state.
Authors
- Kai Wang (ORCID: https://orcid.org/0009-0000-5018-9305)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22739987
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint