Controllable flow fields for high-fidelity virtual try-on with diffusion models

Abstract Virtual try-on is a challenging task that seeks to create realistic images of individuals wearing specific garments and has significant applications in e-commerce and virtual fashion. Existing approaches have made considerable advancements but come with large memory requirements and fail to preserve subtle garment details. To improve on these limitations, we present a new end-to-end diffusion-based framework that allows for more efficient and controllable detail generation. The key component of our approach is the COFFLOW module that extracts precise spatial correspondence from the target garment to the person image and incorporate this into the diffusion process as guidance for image synthesis. We also introduce a blending module that facilitates better visual quality in unmasked areas like face and limbs. We train our model freezing the pretrained diffusion U-Net and updating only the COFFLOW module, leading to a significantly lower cost during training. We evaluate our method on two standard benchmarks, VITON-HD and DressCode. Quantitatively and qualitatively, we show that our method significantly improves upon the prior diffusion-based methods, with better garment fitting, texture preservation, and realistic pose fitting, all while requiring lower computational resources.

Authors

Institutions

Publication Details

Journal
Multimedia Systems
Published
2026-09-21
DOI
https://doi.org/10.1007/s00530-026-02660-9
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Controllable flow fields for high-fidelity virtual try-on with diffusion models

Lai-Man Po, Yu Xue, Xuyuan Xu, Wing-Yin Yu et al.
Multimedia Systems
Generative Adversarial Networks and Image Synthesis
article

Controllable flow fields for high-fidelity virtual try-on with diffusion models

Lai-Man Po, Yu Xue, Xuyuan Xu, Wing-Yin Yu, Yuyang Liu, Kun Li, Zeyu Jiang, Yexin Wang, Haoxuan Wu
article en

Abstract

Abstract Virtual try-on is a challenging task that seeks to create realistic images of individuals wearing specific garments and has significant applications in e-commerce and virtual fashion. Existing approaches have made considerable advancements but come with large memory requirements and fail to preserve subtle garment details. To improve on these limitations, we present a new end-to-end diffusion-based framework that allows for more efficient and controllable detail generation. The key component of our approach is the COFFLOW module that extracts precise spatial correspondence from the target garment to the person image and incorporate this into the diffusion process as guidance for image synthesis. We also introduce a blending module that facilitates better visual quality in unmasked areas like face and limbs. We train our model freezing the pretrained diffusion U-Net and updating only the COFFLOW module, leading to a significantly lower cost during training. We evaluate our method on two standard benchmarks, VITON-HD and DressCode. Quantitatively and qualitatively, we show that our method significantly improves upon the prior diffusion-based methods, with better garment fitting, texture preservation, and realistic pose fitting, all while requiring lower computational resources.

Multimedia SystemsVol. 32(9)
City University of Hong Kong (HK)
Openalex Percentile: Top 13%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Controllable flow fields for high-fidelity virtual try-on with diffusion models — Lai-Man Po, Yu Xue, et al. · Multimedia Systems (2026) | TGRS Research Map | TGRS