Same Local Operation, Different Execution Regime: Conditional Reachability under Multi-Step Task Context

A failed multi-step answer does not uniquely identify the capability that failed. We used Kaiying Reachable Capability Mapping (KRCM) to test six frozen open-weight language models on deterministic addition, symbolic-state, and sequence-rewrite tasks with exact intermediate states. After two earlier estimands were invalidated by adversarial review, we froze a final five-condition, step-4-only experiment in which the target operation and current state were held fixed while the surrounding trace context was manipulated. The final experiment contained 6,282 calls. Relative to a single supplied current state, adding a coherent preceding trajectory changed local reachability in opposite directions across task families: addition improved in five of six models (median +6.5 percentage points), sequence rewriting declined in four of six (median −9.3 points), and symbolic transition performance declined in five of six (median −12.3 points). Under a robustness Holm correction across all 36 primary model×family tests, three trace-presentation cells remained significant: Qwen2.5 ADD (+32.4 points), Qwen2.5 SEQ (−24.7 points), and Qwen3 SYM (−47.5 points). The Qwen3 SYM effect remained interpretable despite a 95.1% single-state baseline because the high reference rate does not constrain the observed negative direction. A value-matched shuffled-trace control showed no stable coherence effect across models. State-referent evidence was strongest in SYM: among model cells with enough errors to estimate the contrast, four of five showed positive observed-minus-null excess, with clear effects in Mistral, Qwen2.5, and Qwen3. In SEQ, a symbol-preserving null retained a pooled excess (24.2% observed versus 12.0% analytic chance among valid-permutation errors), but model-level effects were heterogeneous, with positive excess in four of six models. By contrast, applying the wrong listed operation to the current state was not more frequent among TRACE errors than among SINGLE errors. The supported interpretation is therefore narrower than a generic binding deficit: competing visible state referents can change the reachability of an otherwise executable local transformation in some model-domain regimes. The strongest conclusion is not a universal composition deficit, a working-memory mechanism, or a causal cost of online generation. Local transformation capability is conditionally reachable: the same deterministic operation can become easier or harder when the surrounding execution regime changes.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22740400
Primary Topic
Neurobiology of Language and Bilingualism
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Same Local Operation, Different Execution Regime: Conditional Reachability under Multi-Step Task Context

Kai Wang
Zenodo (CERN European Organization for Nuclear Research)
Neurobiology of Language and Bilingualism
preprint

Same Local Operation, Different Execution Regime: Conditional Reachability under Multi-Step Task Context

Kai Wang
preprint en

Abstract

A failed multi-step answer does not uniquely identify the capability that failed. We used Kaiying Reachable Capability Mapping (KRCM) to test six frozen open-weight language models on deterministic addition, symbolic-state, and sequence-rewrite tasks with exact intermediate states. After two earlier estimands were invalidated by adversarial review, we froze a final five-condition, step-4-only experiment in which the target operation and current state were held fixed while the surrounding trace context was manipulated. The final experiment contained 6,282 calls. Relative to a single supplied current state, adding a coherent preceding trajectory changed local reachability in opposite directions across task families: addition improved in five of six models (median +6.5 percentage points), sequence rewriting declined in four of six (median −9.3 points), and symbolic transition performance declined in five of six (median −12.3 points). Under a robustness Holm correction across all 36 primary model×family tests, three trace-presentation cells remained significant: Qwen2.5 ADD (+32.4 points), Qwen2.5 SEQ (−24.7 points), and Qwen3 SYM (−47.5 points). The Qwen3 SYM effect remained interpretable despite a 95.1% single-state baseline because the high reference rate does not constrain the observed negative direction. A value-matched shuffled-trace control showed no stable coherence effect across models. State-referent evidence was strongest in SYM: among model cells with enough errors to estimate the contrast, four of five showed positive observed-minus-null excess, with clear effects in Mistral, Qwen2.5, and Qwen3. In SEQ, a symbol-preserving null retained a pooled excess (24.2% observed versus 12.0% analytic chance among valid-permutation errors), but model-level effects were heterogeneous, with positive excess in four of six models. By contrast, applying the wrong listed operation to the current state was not more frequent among TRACE errors than among SINGLE errors. The supported interpretation is therefore narrower than a generic binding deficit: competing visible state referents can change the reachability of an otherwise executable local transformation in some model-domain regimes. The strongest conclusion is not a universal composition deficit, a working-memory mechanism, or a causal cost of online generation. Local transformation capability is conditionally reachable: the same deterministic operation can become easier or harder when the surrounding execution regime changes.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Neurobiology of Language and Bilingualism
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.