Sanitizing Music Listening Histories at Scale: A Reproducible Quality Audit of MLHD+
Music recommendation research depends on large-scale datasets of timestamped listening events, yet such datasets are rarely subjected to rigorous, published quality audits. We present a reproducible sanitization and audit pipeline applied to MLHD+ (Music Listening Histories Dataset Plus), the MetaBrainz Foundation's 2023 re-release of the original 27-billion-log MLHD. Auditing one partial shard (384.0 million raw events) and one complete shard (1.30 billion raw events), both covering the same 36,970 users, we document the semantic distinction between MLHD+'s two tiers: partial shards contain only events lacking a resolved recording identifier, a structural property we trace to the official importer source and confirm empirically (zero percent recording-identifier coverage in the partial tier, 100 percent in the complete tier). Our pipeline—schema validation with an itemized rejection taxonomy, event-key deduplication, pseudonymization, and noise flagging rather than deletion—retains 71.9 percent of partial-shard events and 99.2 percent of complete-shard events with full recording-identifier coverage. A failed canonicalization attempt (zero matches across 176 million lookups against the official MessyBrainz mapping tables) is explained structurally and published as a negative result. A cross-shard census extends the structural findings to all 16 complete shards. We release the entire pipeline as deterministic, checksum-verified, versioned kernels under Apache 2.0 and offer the audit as a template for the dataset transparency that listening-history papers have historically omitted.
Authors
- Shuvi (ORCID: https://orcid.org/0000-0002-0628-4361)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-09
- DOI
- https://doi.org/10.5281/zenodo.22338292
- Primary Topic
- Music and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00