Design and implementation of an end to end music audio generation system based on diffusion probabilistic model
Existing music audio generation methods face challenges such as low generation efficiency, a pronounced trade-off between modeling granularity and computational complexity, and incomplete multimodal conditional control mechanisms. To address these issues, this paper proposes an end-to-end music audio generation system based on a diffusion probabilistic model. The system adopts a two-stage “compress-then-generate” architecture. A pre-trained EnCodec neural audio codec compresses the raw audio into low-frame-rate latent vector sequences, and a continuous expansion mechanism is introduced to preserve quantization residual information, thereby enhancing the continuous representational capacity of the latent space. For the noise prediction network, a hybrid architecture combining U-Net and Transformer is constructed, where residual convolutional layers extract local detail features and Transformer layers establish global temporal dependencies, enabling effective modeling of long-range musical structures. In terms of conditional control, a multimodal conditioning mechanism supporting text, melodic reference audio, and rhythmic trajectories is designed, employing classifier-free guidance and a hierarchical control strategy to achieve fine-grained, multidimensional generative control. To improve inference efficiency, the DDIM skip-step sampling and a progressive generation strategy are introduced, significantly reducing the number of sampling steps while enhancing the progressive refinement of musical structure. Experimental results demonstrate that the proposed system outperforms several mainstream audio generation methods across dimensions such as generation quality, conditional controllability, and inference efficiency, with notable improvements in objective metrics including FAD and PESQ, validating the effectiveness of the proposed architecture in efficient generation and precise control. This study provides a novel technical pathway for diffusion-based generation systems targeting complex musical structures.
Authors
- Rui Luo (ORCID: https://orcid.org/0000-0002-4300-121X)
- Yang Yang
- Rong Zhang
Institutions
- Ningxia University (CN)
- Yinchuan First People's Hospital (CN)
- Ningxia Seismological Bureau (CN)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1007/s44163-026-02280-2
- Primary Topic
- Music Technology and Sound Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00