Generative models for interactive content creation and understanding in smart imaging

Generative diffusion models have achieved substantial progress in image content creation, yet practical interactive generation remains challenging when heterogeneous conditions must be jointly interpreted, fine-grained attributes must be controlled independently, and users need to iteratively refine generated results. This paper proposes a Multimodal Interactive Generation Framework (MIGF) that integrates three task-oriented components: Adaptive Condition Fusion (ACF), Semantic Decoupling (SDM), and an Interactive Feedback Mechanism (IFM). ACF addresses sample-dependent reliability differences among text, reference-image, and structured conditions by predicting modality-specific fusion weights for each input. SDM uses contrastive supervision to organize latent representations into content- and style-related subspaces, facilitating more independent attribute control. IFM integrates local modification, attribute adjustment, and example guidance into a unified progressive refinement process. We construct a multimodal dataset containing 50,000 high-resolution images covering natural scenes, artistic works, design patterns, and architecture. Under matched evaluation settings, MIGF reduces FID by 23.7%, improves CLIP Score by 18.6%, and reduces LPIPS by 12.7% relative to the strongest self-run baseline, ControlNet. Ablation experiments show consistent contributions from the three components. The combination of text, image, and sketch conditions achieves a Control Precision of 0.876, corresponding to an 11.9% improvement over the strongest single-modality configuration. In addition, the fast inference mode reduces generation time to 1.9 s per image. These results indicate that the proposed framework provides a practical approach to multimodal and interactively controllable image generation, while its remaining limitations in robustness, computational cost, and cross-domain generalization motivate further investigation.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-05
DOI
https://doi.org/10.1007/s44163-026-02418-2
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Generative models for interactive content creation and understanding in smart imaging

Di Wu, Shuai Bai, Hailong Li
Discover Artificial Intelligence
Generative Adversarial Networks and Image Synthesis
article

Generative models for interactive content creation and understanding in smart imaging

Di Wu, Shuai Bai, Hailong Li
article en

Abstract

Generative diffusion models have achieved substantial progress in image content creation, yet practical interactive generation remains challenging when heterogeneous conditions must be jointly interpreted, fine-grained attributes must be controlled independently, and users need to iteratively refine generated results. This paper proposes a Multimodal Interactive Generation Framework (MIGF) that integrates three task-oriented components: Adaptive Condition Fusion (ACF), Semantic Decoupling (SDM), and an Interactive Feedback Mechanism (IFM). ACF addresses sample-dependent reliability differences among text, reference-image, and structured conditions by predicting modality-specific fusion weights for each input. SDM uses contrastive supervision to organize latent representations into content- and style-related subspaces, facilitating more independent attribute control. IFM integrates local modification, attribute adjustment, and example guidance into a unified progressive refinement process. We construct a multimodal dataset containing 50,000 high-resolution images covering natural scenes, artistic works, design patterns, and architecture. Under matched evaluation settings, MIGF reduces FID by 23.7%, improves CLIP Score by 18.6%, and reduces LPIPS by 12.7% relative to the strongest self-run baseline, ControlNet. Ablation experiments show consistent contributions from the three components. The combination of text, image, and sketch conditions achieves a Control Precision of 0.876, corresponding to an 11.9% improvement over the strongest single-modality configuration. In addition, the fast inference mode reduces generation time to 1.9 s per image. These results indicate that the proposed framework provides a practical approach to multimodal and interactively controllable image generation, while its remaining limitations in robustness, computational cost, and cross-domain generalization motivate further investigation.

Discover Artificial IntelligenceVol. 6(1)
Shandong Huayu University of Technology (CN)
Openalex Percentile: Top 14%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.