Auditing Coding-Agent Rule Updates under Selective Feedback

Persistent coding agents can compile corrections into rules, but selective feedback and changed policies complicate admission. We implement version-scoped finite-pool acceptance, rejection and deferral with executable retirement. Constructed studies show that historical certification need not cover future variants and useful conservative admission approaches a census: four selectors admit no eligible holdout candidate at 20 of 32 labels. Statistical finite-population alternatives preserve this cost. Small external-model studies diagnose stale controls and connect proposals to calibration without establishing superiority over textual memory. A fixed purposive corpus from ten public repositories contains 40 instruction-file transitions and 44 AI-assisted analysis groups, including two explicit local contract replacements; it does not estimate policy prevalence or deployed retirement. A separate pilot runs the real OpenAI Python SDK with two explicitly artificial cap faults. All twelve workers complete and restore the reference, while all four valid matcher proposals are rejected on an authored snippet census. Direct guards never fire and admitted guards never activate, so identical functional outcomes show no admission benefit. These results distinguish documented replacement, finite-pool compatibility and future task behavior. They support reproducible diagnoses, not recursive self-improvement, independent human policy judgments or a general safety guarantee. Exploratory working paper v0.8, first made public on GitHub on 8 October 2026. Not peer reviewed or accepted by a journal. The record mirrors the six files of the exact GitHub v0.8 release at commit 28435e716f8d6453662beb75a6c94e262cfd5f58. The supporting artifact includes constructed audits and model studies, a fixed purposive instruction-file corpus from ten repositories (40 file transitions, 44 AI-assisted analysis groups), and twelve repair workers on two explicitly artificial faults in one pinned real SDK. All four valid matcher proposals were rejected; direct guards never fired and admitted guards never activated. No admission benefit, deployed retirement, independent human policy judgments, multi-repository repair superiority, recursive self-improvement or general safety guarantee is established. The archive preserves third-party licences and provenance; it does not grant a new blanket licence to original or third-party material. The generic manuscript.tex companion is not the venue-formatted submission source and does not reproduce the three plots included in the canonical PDF. The canonical frozen paper itself accurately records author-fact verification as pending at the time of the GitHub release; any subsequently confirmed author declarations must be recorded separately and do not silently modify v0.8. GitHub release: https://github.com/Linxiushen/coding-agent-rule-audit/releases/tag/v0.8Repository: https://github.com/Linxiushen/coding-agent-rule-auditThird-party notices: https://github.com/Linxiushen/coding-agent-rule-audit/blob/v0.8/THIRD_PARTY_NOTICES.md

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23234125
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Auditing Coding-Agent Rule Updates under Selective Feedback

Xulin Chen
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
article

Auditing Coding-Agent Rule Updates under Selective Feedback

Xulin Chen
article en

Abstract

Persistent coding agents can compile corrections into rules, but selective feedback and changed policies complicate admission. We implement version-scoped finite-pool acceptance, rejection and deferral with executable retirement. Constructed studies show that historical certification need not cover future variants and useful conservative admission approaches a census: four selectors admit no eligible holdout candidate at 20 of 32 labels. Statistical finite-population alternatives preserve this cost. Small external-model studies diagnose stale controls and connect proposals to calibration without establishing superiority over textual memory. A fixed purposive corpus from ten public repositories contains 40 instruction-file transitions and 44 AI-assisted analysis groups, including two explicit local contract replacements; it does not estimate policy prevalence or deployed retirement. A separate pilot runs the real OpenAI Python SDK with two explicitly artificial cap faults. All twelve workers complete and restore the reference, while all four valid matcher proposals are rejected on an authored snippet census. Direct guards never fire and admitted guards never activate, so identical functional outcomes show no admission benefit. These results distinguish documented replacement, finite-pool compatibility and future task behavior. They support reproducible diagnoses, not recursive self-improvement, independent human policy judgments or a general safety guarantee. Exploratory working paper v0.8, first made public on GitHub on 8 October 2026. Not peer reviewed or accepted by a journal. The record mirrors the six files of the exact GitHub v0.8 release at commit 28435e716f8d6453662beb75a6c94e262cfd5f58. The supporting artifact includes constructed audits and model studies, a fixed purposive instruction-file corpus from ten repositories (40 file transitions, 44 AI-assisted analysis groups), and twelve repair workers on two explicitly artificial faults in one pinned real SDK. All four valid matcher proposals were rejected; direct guards never fired and admitted guards never activated. No admission benefit, deployed retirement, independent human policy judgments, multi-repository repair superiority, recursive self-improvement or general safety guarantee is established. The archive preserves third-party licences and provenance; it does not grant a new blanket licence to original or third-party material. The generic manuscript.tex companion is not the venue-formatted submission source and does not reproduce the three plots included in the canonical PDF. The canonical frozen paper itself accurately records author-fact verification as pending at the time of the GitHub release; any subsequently confirmed author declarations must be recorded separately and do not silently modify v0.8. GitHub release: https://github.com/Linxiushen/coding-agent-rule-audit/releases/tag/v0.8Repository: https://github.com/Linxiushen/coding-agent-rule-auditThird-party notices: https://github.com/Linxiushen/coding-agent-rule-audit/blob/v0.8/THIRD_PARTY_NOTICES.md

Zenodo (CERN European Organization for Nuclear Research)
Lynx (Italy) (IT)
Openalex Percentile: Top 5%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.