DTMG-Net: diffusion model with MoE and multi-scale dilated attention for CRC tongue image generative augmentation

Deep-learning-based image generation and classification have significantly advanced computer-aided diagnosis. However, the scarcity of high-quality labeled tongue image datasets for colorectal cancer (CRC) diagnosing tasks severely limits model generalization. While conventional data augmentation fails to mimic complex clinically relevant pathological variations, generative data augmentation offers a promising alternative. This work proposes DTMG-Net, an improved DDIM generative augmentation framework. We embed the DeepSeekMoE module into the denoising U-Net to realize adaptive modeling across diffusion timesteps of multi-scale pathological features. A multi-scale dilated attention (MSDA) block is introduced in the bottleneck to jointly capture global tongue structures and subtle local lesion textures. Moreover, a joint loss combining standard noise estimation and intra-sample diversity regularization is designed to alleviate mode collapse and expand the visual diversity of synthetic tongue samples. We adopt FID and IS to quantitatively evaluate the quality of generated images, and perform downstream classification experiments on four mainstream backbone networks to verify the effectiveness of synthetic data. Compared with VAE, DCGAN, PNDM and original DDIM, DTMG-Net achieves lower FID values of 73.83 and 58.99 for CRC and HC tongue samples respectively. Although these values remain relatively high in absolute terms, our method attains the minimal distribution discrepancy among all compared generative models under the small-sample tongue image setting. Ablation experiments indicate that DeepSeekMoE and MSDA both contribute to improved generation performance, and the proposed diversity constraint elevates IS without obvious FID degradation. In downstream classification tasks, training sets augmented with DTMG-Net synthetic images achieve generally higher numerical AUC, F1-score and Accuracy on WideResNet, ResNet50, MedMamba and ViT among the compared augmentation strategies. The proposed DTMG-Net can generate high-diversity, high-fidelity tongue samples to effectively expand small-scale datasets. The intra-sample diversity loss balances reconstruction quality and visual richness of generated content. Without complex preprocessing operations, this diffusion-based generative augmentation scheme achieves generally the highest numerical classification performance among the compared strategies on various classification backbones, supporting the potential of combining improved MoE and multi-scale attention modules for small-sample tongue image augmentation tasks.

Authors

Institutions

Publication Details

Journal
BMC Medical Imaging
Published
2026-09-08
DOI
https://doi.org/10.1186/s12880-026-02755-9
Primary Topic
Voice and Speech Disorders
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

DTMG-Net: diffusion model with MoE and multi-scale dilated attention for CRC tongue image generative augmentation

Decao Niu, Dabiao Wang, Chongyang Wang, Lanlan Li et al.
BMC Medical Imaging
Voice and Speech Disorders
article

DTMG-Net: diffusion model with MoE and multi-scale dilated attention for CRC tongue image generative augmentation

Decao Niu, Dabiao Wang, Chongyang Wang, Lanlan Li, Yanhao Ren, Liqun Lin, Juan Li, Weifeng Liu, Ziyue Wang, Xuan Yang, Yu Zeng
article en

Abstract

Deep-learning-based image generation and classification have significantly advanced computer-aided diagnosis. However, the scarcity of high-quality labeled tongue image datasets for colorectal cancer (CRC) diagnosing tasks severely limits model generalization. While conventional data augmentation fails to mimic complex clinically relevant pathological variations, generative data augmentation offers a promising alternative. This work proposes DTMG-Net, an improved DDIM generative augmentation framework. We embed the DeepSeekMoE module into the denoising U-Net to realize adaptive modeling across diffusion timesteps of multi-scale pathological features. A multi-scale dilated attention (MSDA) block is introduced in the bottleneck to jointly capture global tongue structures and subtle local lesion textures. Moreover, a joint loss combining standard noise estimation and intra-sample diversity regularization is designed to alleviate mode collapse and expand the visual diversity of synthetic tongue samples. We adopt FID and IS to quantitatively evaluate the quality of generated images, and perform downstream classification experiments on four mainstream backbone networks to verify the effectiveness of synthetic data. Compared with VAE, DCGAN, PNDM and original DDIM, DTMG-Net achieves lower FID values of 73.83 and 58.99 for CRC and HC tongue samples respectively. Although these values remain relatively high in absolute terms, our method attains the minimal distribution discrepancy among all compared generative models under the small-sample tongue image setting. Ablation experiments indicate that DeepSeekMoE and MSDA both contribute to improved generation performance, and the proposed diversity constraint elevates IS without obvious FID degradation. In downstream classification tasks, training sets augmented with DTMG-Net synthetic images achieve generally higher numerical AUC, F1-score and Accuracy on WideResNet, ResNet50, MedMamba and ViT among the compared augmentation strategies. The proposed DTMG-Net can generate high-diversity, high-fidelity tongue samples to effectively expand small-scale datasets. The intra-sample diversity loss balances reconstruction quality and visual richness of generated content. Without complex preprocessing operations, this diffusion-based generative augmentation scheme achieves generally the highest numerical classification performance among the compared strategies on various classification backbones, supporting the potential of combining improved MoE and multi-scale attention modules for small-sample tongue image augmentation tasks.

BMC Medical Imaging
Sun Yat-sen University (CN), Sixth Affiliated Hospital of Sun Yat-sen University (CN), Guangdong Provincial People's Hospital (CN), Fuzhou University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 11%
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.