How Many Heads, and Ranked How: Two Unreported Choices That Determine a Circuit
A circuit found by ranking attention heads and keeping the top k is specified by two choices - the value of k, and the rule used to rank - and neither is typically reported. Across ten distinct discovery configurations in Llama-3.2-3B and Pythia-1.4B, sweeping k from 1 to 50 with a paired bootstrap over examples, ten heads (the value inherited through most of this literature) reaches faithfulness F >= 0.8 in two of nine interpretable configurations and F >= 0.95 in none. F rises monotonically and passes 1.0 in eight of nine configurations, so any faithfulness threshold is satisfiable by adding heads. The ranking rule matters comparably: three defensible rules built from the same stored measurements disagree, with a median spread of 0.176 and a maximum of 0.837 at k=10, and no rule dominates - patching is the worse choice for Pythia IOI and the better one for Llama IOI. Completeness, measured separately, does not track faithfulness (rank correlation +0.200 at k=20): the most complete circuit found has F=0.901, and every configuration whose F exceeds 1.0 retains a completeness gap of 0.097-0.221. A bootstrap-based non-degeneracy gate additionally catches a configuration that a conventional point-estimate check would certify, whose F was the largest and whose knee was the earliest in the sweep - both artifacts of a denominator that changes sign under resampling. All runs on a single Apple M1 Max.
Authors
- J. Melton
Institutions
- American Standard (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-11
- DOI
- https://doi.org/10.5281/zenodo.22700818
- Primary Topic
- Cognitive and developmental aspects of mathematical skills
- Type
- preprint