Controllable flow fields for high-fidelity virtual try-on with diffusion models
Abstract Virtual try-on is a challenging task that seeks to create realistic images of individuals wearing specific garments and has significant applications in e-commerce and virtual fashion. Existing approaches have made considerable advancements but come with large memory requirements and fail to preserve subtle garment details. To improve on these limitations, we present a new end-to-end diffusion-based framework that allows for more efficient and controllable detail generation. The key component of our approach is the COFFLOW module that extracts precise spatial correspondence from the target garment to the person image and incorporate this into the diffusion process as guidance for image synthesis. We also introduce a blending module that facilitates better visual quality in unmasked areas like face and limbs. We train our model freezing the pretrained diffusion U-Net and updating only the COFFLOW module, leading to a significantly lower cost during training. We evaluate our method on two standard benchmarks, VITON-HD and DressCode. Quantitatively and qualitatively, we show that our method significantly improves upon the prior diffusion-based methods, with better garment fitting, texture preservation, and realistic pose fitting, all while requiring lower computational resources.
Authors
- Lai-Man Po (ORCID: https://orcid.org/0000-0002-5185-1492)
- Yu Xue (ORCID: https://orcid.org/0009-0006-3961-890X)
- Xuyuan Xu (ORCID: https://orcid.org/0009-0000-6609-1046)
- Wing-Yin Yu (ORCID: https://orcid.org/0000-0002-9559-1055)
- Yuyang Liu (ORCID: https://orcid.org/0000-0003-0418-4989)
- Kun Li (ORCID: https://orcid.org/0009-0006-2944-3501)
- Zeyu Jiang (ORCID: https://orcid.org/0009-0009-9011-2877)
- Yexin Wang (ORCID: https://orcid.org/0000-0003-4873-7127)
- Haoxuan Wu (ORCID: https://orcid.org/0009-0009-1710-4408)
Institutions
- City University of Hong Kong (HK)
Publication Details
- Journal
- Multimedia Systems
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1007/s00530-026-02660-9
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00