Sanitizing Music Listening Histories at Scale: A Reproducible Quality Audit of MLHD+

Music recommendation research depends on large-scale datasets of timestamped listening events, yet such datasets are rarely subjected to rigorous, published quality audits. We present a reproducible sanitization and audit pipeline applied to MLHD+ (Music Listening Histories Dataset Plus), the MetaBrainz Foundation's 2023 re-release of the original 27-billion-log MLHD. Auditing one partial shard (384.0 million raw events) and one complete shard (1.30 billion raw events), both covering the same 36,970 users, we document the semantic distinction between MLHD+'s two tiers: partial shards contain only events lacking a resolved recording identifier, a structural property we trace to the official importer source and confirm empirically (zero percent recording-identifier coverage in the partial tier, 100 percent in the complete tier). Our pipeline—schema validation with an itemized rejection taxonomy, event-key deduplication, pseudonymization, and noise flagging rather than deletion—retains 71.9 percent of partial-shard events and 99.2 percent of complete-shard events with full recording-identifier coverage. A failed canonicalization attempt (zero matches across 176 million lookups against the official MessyBrainz mapping tables) is explained structurally and published as a negative result. A cross-shard census extends the structural findings to all 16 complete shards. We release the entire pipeline as deterministic, checksum-verified, versioned kernels under Apache 2.0 and offer the audit as a template for the dataset transparency that listening-history papers have historically omitted.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-09
DOI
https://doi.org/10.5281/zenodo.22338292
Primary Topic
Music and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Sanitizing Music Listening Histories at Scale: A Reproducible Quality Audit of MLHD+

Shuvi
Zenodo (CERN European Organization for Nuclear Research)
Music and Audio Processing
article

Sanitizing Music Listening Histories at Scale: A Reproducible Quality Audit of MLHD+

Shuvi
article en

Abstract

Music recommendation research depends on large-scale datasets of timestamped listening events, yet such datasets are rarely subjected to rigorous, published quality audits. We present a reproducible sanitization and audit pipeline applied to MLHD+ (Music Listening Histories Dataset Plus), the MetaBrainz Foundation's 2023 re-release of the original 27-billion-log MLHD. Auditing one partial shard (384.0 million raw events) and one complete shard (1.30 billion raw events), both covering the same 36,970 users, we document the semantic distinction between MLHD+'s two tiers: partial shards contain only events lacking a resolved recording identifier, a structural property we trace to the official importer source and confirm empirically (zero percent recording-identifier coverage in the partial tier, 100 percent in the complete tier). Our pipeline—schema validation with an itemized rejection taxonomy, event-key deduplication, pseudonymization, and noise flagging rather than deletion—retains 71.9 percent of partial-shard events and 99.2 percent of complete-shard events with full recording-identifier coverage. A failed canonicalization attempt (zero matches across 176 million lookups against the official MessyBrainz mapping tables) is explained structurally and published as a negative result. A cross-shard census extends the structural findings to all 16 complete shards. We release the entire pipeline as deterministic, checksum-verified, versioned kernels under Apache 2.0 and offer the audit as a template for the dataset transparency that listening-history papers have historically omitted.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 9%
Music and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Sanitizing Music Listening Histories at Scale: A Reproducible Quality Audit of MLHD+ — Shuvi · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS