Calibrating Circuit Discovery: Algorithm Choice Against Rule Choice
Two published comparisons of circuit discovery algorithms reach opposite orderings of attribution patching against ACDC, and neither has a yardstick for how much disagreement between methods should be expected. This measures the within-method term first - changing only the rule used to rank candidate heads inside top-k selection moves faithfulness by up to 0.837 at fixed circuit size - then measures the corresponding between-method term on the same harness. Three selection procedures (marginal top-k, greedy forward selection, and reverse-order ACDC-style threshold pruning) are compared at matched circuit size across nine configurations in Llama-3.2-3B and Pythia-1.4B. An uncontrolled first pass, in which one fixed pruning threshold produced circuits from 3 to 33 heads, found the within-method term larger (median ratio 0.50). A second pass with circuit size matched near 15 heads per configuration reverses this: the between-method term is larger in six of nine configurations (median ratio 2.72, range 0.36-13.81), and eight of nine show statistically significant separation between algorithms at 95% confidence. The three exceptions are explained rather than averaged over: two reproduce an already-documented ranking-rule pathology on Pythia IOI, and one is ordinary rule disagreement. No attribution-patching arm is implemented, so this does not adjudicate the contested literature directly, but supplies the yardstick against which it should be read. All runs on a single Apple M1 Max.
Authors
- J. Melton
Institutions
- American Standard (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-11
- DOI
- https://doi.org/10.5281/zenodo.22700816
- Primary Topic
- VLSI and FPGA Design Techniques
- Type
- preprint