Embodied Musical Gesture Analysis from Audio Representations: A Review of Datasets, Audio Descriptors, and Deep Learning Approaches

Machine learning approaches have enabled increasingly powerful analysis of musical audio, allowing computational models to capture complex acoustic patterns related to musical structure, performance, and expression. However, whether learned audio representations preserve information related to the bodily actions underlying musical sound production remains an open research question. This review examines how machine-learned audio representations encode information associated with embodied musical processes by connecting perspectives from embodied music cognition, computational music analysis, and multimodal machine learning. It establishes a conceptual framework that organizes embodied musical information according to gesture type, acoustic observability, and the relationship between physical action and sound production, distinguishing sound-producing gestures, expressive movements, individual embodiment, and collective performance processes. The review then surveys datasets for audio–gesture research, including motion-capture recordings, instrument-performance datasets, audiovisual resources, and large-scale audio-only datasets, highlighting challenges related to embodiment fidelity, scalability, annotation, and multimodal alignment. It further examines conventional audio descriptors, deep neural and self-supervised audio representations, cross-modal learning approaches, and evaluation methodologies for assessing embodied information within learned representations. Particular attention is given to gesture–sound ambiguity, limited benchmarking resources, generalization across performers and instruments, and the distinction between gesture-related encoding and acoustic correlation. Current evidence suggests that learned audio representations contain partial and task-dependent traces of embodied musical information, particularly for gestures with strong acoustic coupling, while evidence for higher-level expressive and communicative movements remains limited. Determining whether these representations capture meaningful relationships between physical performance processes and musical sound remains an open challenge requiring multimodal, interpretable, and perceptually grounded evaluation frameworks.

Authors

Institutions

Publication Details

Journal
Arts
Published
2026-09-22
DOI
https://doi.org/10.3390/arts15100217
Primary Topic
Music Technology and Sound Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Embodied Musical Gesture Analysis from Audio Representations: A Review of Datasets, Audio Descriptors, and Deep Learning Approaches

Ioanna Miliaresi
Arts
Music Technology and Sound Studies
article

Embodied Musical Gesture Analysis from Audio Representations: A Review of Datasets, Audio Descriptors, and Deep Learning Approaches

Ioanna Miliaresi
article en

Abstract

Machine learning approaches have enabled increasingly powerful analysis of musical audio, allowing computational models to capture complex acoustic patterns related to musical structure, performance, and expression. However, whether learned audio representations preserve information related to the bodily actions underlying musical sound production remains an open research question. This review examines how machine-learned audio representations encode information associated with embodied musical processes by connecting perspectives from embodied music cognition, computational music analysis, and multimodal machine learning. It establishes a conceptual framework that organizes embodied musical information according to gesture type, acoustic observability, and the relationship between physical action and sound production, distinguishing sound-producing gestures, expressive movements, individual embodiment, and collective performance processes. The review then surveys datasets for audio–gesture research, including motion-capture recordings, instrument-performance datasets, audiovisual resources, and large-scale audio-only datasets, highlighting challenges related to embodiment fidelity, scalability, annotation, and multimodal alignment. It further examines conventional audio descriptors, deep neural and self-supervised audio representations, cross-modal learning approaches, and evaluation methodologies for assessing embodied information within learned representations. Particular attention is given to gesture–sound ambiguity, limited benchmarking resources, generalization across performers and instruments, and the distinction between gesture-related encoding and acoustic correlation. Current evidence suggests that learned audio representations contain partial and task-dependent traces of embodied musical information, particularly for gestures with strong acoustic coupling, while evidence for higher-level expressive and communicative movements remains limited. Determining whether these representations capture meaningful relationships between physical performance processes and musical sound remains an open challenge requiring multimodal, interpretable, and perceptually grounded evaluation frameworks.

ArtsVol. 15(10)
Ionian University (GR)
Quality Education
Openalex Percentile: Top 13%
Music Technology and Sound Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.