Layout-Conditioned Flow Matching with Mask-Guided Region Pooling For Anime Face Generation
We train an anime-face generator from scratch on 21,551 images using flow matching, with a network that predicts the clean image directly rather than a velocity. The training loss is the plain flow-matching objective with no auxiliary pixel-space terms: an ablation over five such losses (edge, line, palette) found each one neutral or harmful, consistent with the fact that such losses are largely irreducible at the objective's optimum. We instead fix a specific failure mode, mismatched left/right iris colors, with an architectural change: conditioning on a coarse layout (face, eye, and mouth regions from detected landmarks), plus a zero-initialized pooling layer that shares features within each region, cuts the mismatch rate from 35% to 3.5%, near the real data's own 4 to 6%. At sampling time, layouts are drawn from a fitted prior over the landmarks, so no real image is used, and this model reaches FID 28.8 on prior layouts and 28.3 on real layouts. We further apply two existing techniques, autoguidance (Karras et al.), which improves FID to 21.2 at no retraining cost, and teacher distillation into a 1 to 2 step sampler for CPU-only inference; the same approach transfers to CelebAMask-HQ with real label maps, reaching FID 27.3 and 18.6 with autoguidance.
Authors
- Nikhil Raghavendra
Institutions
- PES University (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22876650
- Primary Topic
- Face recognition and analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00