Embodied Musical Gesture Analysis from Audio Representations: A Review of Datasets, Audio Descriptors, and Deep Learning Approaches
Machine learning approaches have enabled increasingly powerful analysis of musical audio, allowing computational models to capture complex acoustic patterns related to musical structure, performance, and expression. However, whether learned audio representations preserve information related to the bodily actions underlying musical sound production remains an open research question. This review examines how machine-learned audio representations encode information associated with embodied musical processes by connecting perspectives from embodied music cognition, computational music analysis, and multimodal machine learning. It establishes a conceptual framework that organizes embodied musical information according to gesture type, acoustic observability, and the relationship between physical action and sound production, distinguishing sound-producing gestures, expressive movements, individual embodiment, and collective performance processes. The review then surveys datasets for audio–gesture research, including motion-capture recordings, instrument-performance datasets, audiovisual resources, and large-scale audio-only datasets, highlighting challenges related to embodiment fidelity, scalability, annotation, and multimodal alignment. It further examines conventional audio descriptors, deep neural and self-supervised audio representations, cross-modal learning approaches, and evaluation methodologies for assessing embodied information within learned representations. Particular attention is given to gesture–sound ambiguity, limited benchmarking resources, generalization across performers and instruments, and the distinction between gesture-related encoding and acoustic correlation. Current evidence suggests that learned audio representations contain partial and task-dependent traces of embodied musical information, particularly for gestures with strong acoustic coupling, while evidence for higher-level expressive and communicative movements remains limited. Determining whether these representations capture meaningful relationships between physical performance processes and musical sound remains an open challenge requiring multimodal, interpretable, and perceptually grounded evaluation frameworks.
Authors
- Ioanna Miliaresi (ORCID: https://orcid.org/0000-0002-4461-9206)
Institutions
- Ionian University (GR)
Publication Details
- Journal
- Arts
- Published
- 2026-09-22
- DOI
- https://doi.org/10.3390/arts15100217
- Primary Topic
- Music Technology and Sound Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00