Music Generation Model Based on Multi-Encoder Transformer and Improved GAN

Driven by the advancement of artificial intelligence in music generation and cross-modal composition, achieving coherence of musical themes and accurate matching of background music in videos has become a research hotspot. To this end, this study constructs a music generation model with a multi-encoder Transformer and an improved generative adversarial network. This model combines a basic encoder and a time-series encoder to enhance the capture of musical themes and achieves selective modeling of theme information through an adaptive theme decoder. Simultaneously, it achieves accurate alignment of video motion information with musical rhythm through a spatial channel attention mechanism and multi-scale convolution optimization of the generator and discriminator. Findings indicate that in four music styles, piano, jazz, electronic, and rock, the model’s average ratio of empty bars is 11–13%, the number of used pitch classes is 44–46, and the ratio of qualified notes is 86–89%. Furthermore, in a listening test, the model’s overall subjective evaluation score is 4.36[Formula: see text] ± [Formula: see text]0.25, close to the 4.53[Formula: see text] ± [Formula: see text]0.22 score of real music samples. Therefore, the model demonstrates high accuracy, naturalness, and robustness in multi-style music generation and cross-modal composition, offering an efficient intelligent solution for intelligent composition and human-computer interactive music creation.

Authors

Institutions

Publication Details

Journal
International Journal of Computational Intelligence and Applications
Published
2026-09-29
DOI
https://doi.org/10.1142/s1469026826500537
Primary Topic
Music Technology and Sound Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Music Generation Model Based on Multi-Encoder Transformer and Improved GAN

Ruoqi Wang
International Journal of Computational Intelligence and Applications
Music Technology and Sound Studies
article

Music Generation Model Based on Multi-Encoder Transformer and Improved GAN

Ruoqi Wang
article en

Abstract

Driven by the advancement of artificial intelligence in music generation and cross-modal composition, achieving coherence of musical themes and accurate matching of background music in videos has become a research hotspot. To this end, this study constructs a music generation model with a multi-encoder Transformer and an improved generative adversarial network. This model combines a basic encoder and a time-series encoder to enhance the capture of musical themes and achieves selective modeling of theme information through an adaptive theme decoder. Simultaneously, it achieves accurate alignment of video motion information with musical rhythm through a spatial channel attention mechanism and multi-scale convolution optimization of the generator and discriminator. Findings indicate that in four music styles, piano, jazz, electronic, and rock, the model’s average ratio of empty bars is 11–13%, the number of used pitch classes is 44–46, and the ratio of qualified notes is 86–89%. Furthermore, in a listening test, the model’s overall subjective evaluation score is 4.36[Formula: see text] ± [Formula: see text]0.25, close to the 4.53[Formula: see text] ± [Formula: see text]0.22 score of real music samples. Therefore, the model demonstrates high accuracy, naturalness, and robustness in multi-style music generation and cross-modal composition, offering an efficient intelligent solution for intelligent composition and human-computer interactive music creation.

International Journal of Computational Intelligence and Applications
Weinan Normal University (CN)
Reduced inequalities
Openalex Percentile: Top 14%
Music Technology and Sound Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Music Generation Model Based on Multi-Encoder Transformer and Improved GAN — Ruoqi Wang · International Journal of Computational Intelligence and Applications (2026) | TGRS Research Map | TGRS