Input Calibration and a Specificity Audit of Auditory-Scene-Analysis Cue Scores in Music Foundation Models
Do music foundation models encode auditory-scene-analysis (ASA) cues, the grouping principles behind source perception? A high cue score alone cannot answer this: the ordering may be inherited from the model's input transform or, for timing, arise from onset positions alone. We introduce the Bregman Cue-Sensitivity Score (BCS), a training-free geometric read-out of four ASA cues, with a reference for each explanation: the matched front-end under the same operator, and index-only references for the temporal cue. Across sixteen checkpoints (thirteen in the main map, ten with a matched front-end; MusicGen-S/M/L as an extraction case study), tested fixed transforms match or exceed the highest main-map checkpoint score on three cues, and paired contrasts fall on both sides of the front-end. For temporal proximity, a model-free clock scores above every checkpoint; only 3 of thirteen checkpoints resolvably exceed their shared onset trajectory. Cue probes should therefore report each applicable reference.
Authors
- Von‐Wun Soo (ORCID: https://orcid.org/0000-0002-4810-1244)
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- Chang Gung University (TW)
- National Tsing Hua University (TW)
- North Carolina Exploring Cultural Heritage Online (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22933795
- Primary Topic
- Neuroscience and Music Perception
- Type
- preprint