Cartoon style generation of animated image scenes based on improved GAN and RetNet
To improve the quality and style consistency of cartoon generation in animated image scenes, this study proposes an image cartoon generation model that integrates improved generative adversarial networks and preservation networks. The model achieves collaborative optimization of local details, global structure, and diverse styles in images by introducing a multi-style mapping module, a multi-head discriminator, a Euclidean distance decay attention mechanism, and a thin plate spline geometric consistency module. This study employs two publicly available scene image datasets, setting up four typical simulation testing environments and comparing them with three mainstream image generation methods. The experiment showed that the model performed well in key indicators such as generation quality, processing efficiency, and structural preservation, with a peak signal-to-noise ratio of up to 29.2 dB, the Fréchet Inception Distance was 35.2, and a processing efficiency of 8.0 FPS. In addition, the generated cartoon image had a structural similarity index of 0.89 with the original image, and the running memory usage was controlled within 3.9GB. The image generation model has overcome the style drift and geometric distortion in existing animated cartoonization, providing a high-fidelity and low-cost automated solution for film and television production, game development, and digital content creation.
Authors
- Zhenyao Jin
- Lei Yang (ORCID: https://orcid.org/0009-0003-9356-658X)
Institutions
- Cheongju University (KR)
- Shandong University of Art and Design (CN)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1371/journal.pone.0347014
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00