Same Local Operation, Different Execution Regime: Conditional Reachability under Multi-Step Task Context
A failed multi-step answer does not uniquely identify the capability that failed. We used Kaiying Reachable Capability Mapping (KRCM) to test six frozen open-weight language models on deterministic addition, symbolic-state, and sequence-rewrite tasks with exact intermediate states. After two earlier estimands were invalidated by adversarial review, we froze a final five-condition, step-4-only experiment in which the target operation and current state were held fixed while the surrounding trace context was manipulated. The final experiment contained 6,282 calls. Relative to a single supplied current state, adding a coherent preceding trajectory changed local reachability in opposite directions across task families: addition improved in five of six models (median +6.5 percentage points), sequence rewriting declined in four of six (median −9.3 points), and symbolic transition performance declined in five of six (median −12.3 points). Under a robustness Holm correction across all 36 primary model×family tests, three trace-presentation cells remained significant: Qwen2.5 ADD (+32.4 points), Qwen2.5 SEQ (−24.7 points), and Qwen3 SYM (−47.5 points). The Qwen3 SYM effect remained interpretable despite a 95.1% single-state baseline because the high reference rate does not constrain the observed negative direction. A value-matched shuffled-trace control showed no stable coherence effect across models. State-referent evidence was strongest in SYM: among model cells with enough errors to estimate the contrast, four of five showed positive observed-minus-null excess, with clear effects in Mistral, Qwen2.5, and Qwen3. In SEQ, a symbol-preserving null retained a pooled excess (24.2% observed versus 12.0% analytic chance among valid-permutation errors), but model-level effects were heterogeneous, with positive excess in four of six models. By contrast, applying the wrong listed operation to the current state was not more frequent among TRACE errors than among SINGLE errors. The supported interpretation is therefore narrower than a generic binding deficit: competing visible state referents can change the reachability of an otherwise executable local transformation in some model-domain regimes. The strongest conclusion is not a universal composition deficit, a working-memory mechanism, or a causal cost of online generation. Local transformation capability is conditionally reachable: the same deterministic operation can become easier or harder when the surrounding execution regime changes.
Authors
- Kai Wang (ORCID: https://orcid.org/0009-0000-5018-9305)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22740400
- Primary Topic
- Neurobiology of Language and Bilingualism
- Type
- preprint