Auditing behavioral claims about AI: Two public-data applications
Behavioral claims about AI-mediated platforms often connect an intervention delivered to one actor with responses by another, yet public records may reveal only selected links. We present a prototype, estimand-specific audit that separately evaluates treatment recoverability, assignment-mechanism fidelity, outcome-population observability, counterfactual adequacy, and behavioral-mechanism discriminability. A deterministic rule system converts source-addressable evidence states into coordinate-specific claim profiles; execution is reproducible conditional on coding, not independently validated. Two retrospective applications expose complementary weaknesses. In Wikipedia's 2025 RevertRisk rollout, logs establish interface availability but not reviewer use or contributor awareness, and purposeful batching precludes a causal effect interpretation. For a fixed observable cohort, field-compatible completion yields an attainable descriptive interval of [0.017316, 0.019146] for one binary feedback channel, not a causal identified set. In public files from a registered randomized generative-AI experiment, intended allocation is documented, but repeated identifiers and unresolved first-assignment lineage prevent third-party reconstruction of a person-level intention-to-treat mapping; this does not challenge the original experiment's randomization. Neither application distinguishes discouragement, repair, or other behavioral mechanisms. The audit provides a fail-closed, noncompensating record of the strongest claims licensed by public evidence and the missing links required to strengthen them.
Authors
- Kaibo Tang (ORCID: https://orcid.org/0009-0004-2637-8025)
Institutions
- Osaka University of Economics (JP)
- Toneyama National Hospital (JP)
Publication Details
- Journal
- Social Sciences & Humanities Open
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1016/j.ssaho.2026.103760
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00