Construct Admission for AI Capability Operators: An Adversarial Protocol and a Worked Failure of Unguarded Dissociation
We introduce KCP, a construct-admission protocol for deciding when a proposed AI capability operator should be merged, decomposed, rejected, left unresolved, or provisionally retained. Its contribution is methodological rather than ontological: KCP requires explicit attack conditions before a capability distinction is treated as interpretable. A six-model case study provides a worked failure of unguarded dissociation reasoning. Same-object JUDGE–EXTENT comparisons initially appeared to support a cross-model distinction. Successive audits showed that this inference was not licensed. On the Boolean-CNF paired subset, the probes do not clear the relevant response-floor conditions. On a second string-constraint family, the nominally strongest case (Mistral-Small-24B) is also unresolved: multiplicity removes the paired floor claim, the item-level association is imprecise, and a simple shared-versus-independent solvability sensitivity analysis shows that 𝑛 = 60 has little discriminating power. The observed response policies are also biased rather than uniform guesses, so the toy solve-or-guess calculation cannot identify the true response process. These failures motivate a reusable KCP Admission Checklist: (1) inspect response-policy informativeness; (2) recompute the floor on the exact paired subset; (3) preregister multiplicity handling; (4) quantify response-space guessability 𝑔; (5) register competing response-process models; (6) demonstrate prospective discriminability/power under independent pilot- or prior-based parameter assumptions; and only then (7) interpret coupling or discordance as construct evidence. In a symmetric solve-or-guess toy model, the separation between fully shared and independent solvability is Δ = 𝑠(1 − 𝑠)(1 −𝑔)2, showing that response-space guessability can quadratically attenuate construct discriminability within that model. This connects two otherwise separate failures in the study: the effective collapse of a paired EXTENT contrast to a binary response space and the weak evidential value of a multiple-choice INTEGRATE intervention. Other controls likewise lower claim strength: matched EXTENT generation rejects a uniqueness-specific interpretation, INDUCE is confirmed to mix identification and application, and INTEGRATE remains unresolved after accounting for multiple-choice funnelling. We do not claim that prior LLM-evaluation papers systematically commit the exact error demonstrated here; no systematic practice audit is performed. The narrower claim is that matched discordance alone does not rule out floor, response-policy, guessability, or discriminability explanations, and that these guards should be made explicit when capability distinctions are inferred. The paper therefore ends without a validated operator separation. Its output is a protocol: do not infer a capability distinction from probes that are uninformative or from a design that cannot distinguish the registered alternative response processes.
Authors
- Kai Wang (ORCID: https://orcid.org/0009-0000-5018-9305)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22767899
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint