SMFLIX: A Unified Body–Head Model for Real-Time Expressive Human Mesh Recovery
Expressive human mesh recovery reconstructs the body, hands, and face from a single image. Whole-body models such as SMPL-X offer only coarse control over the face, whereas face models such as FLAME are head-only. Combining them requires maintaining two models at inference and re-annotating training data. Complete eye closure and lip articulation are perceptually important yet difficult to represent, supervise, and evaluate. Existing formulations leave residual gaps or artifacts, and no public dataset, to our knowledge, characterizes eye and lip aperture. We present SMFLIX, a unified body–head model that merges both bases into a single linear block matrix formulation. It eliminates external registration, yields a watertight neck seam, and makes these articulations representable. Parameter compatibility lets us build SMFLIX-Set without full re-annotation, yielding 11M instances with joint whole-body and facial supervision. The dataset includes 536K face-valid test samples with continuous eye and lip aperture ratios, enabling direct evaluation of articulations obscured by mesh-averaged metrics. We train SMFLIX-Net, a one-stage dual-decoder network regressing body and facial parameters separately. Compared with state-of-the-art whole-body, face-only, and combined methods, the network achieves the lowest whole-body, body-part, and aperture errors on SMFLIX-Set. It remains competitive on external benchmarks, runs at 48 FPS on a consumer GPU, and supports a live webcam application.
Authors
- 오윤성
- Young-Woon Cha (ORCID: https://orcid.org/0000-0003-2234-7729)
- Seokhyeon Heo
Institutions
- Konkuk University (KR)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-10
- DOI
- https://doi.org/10.3390/math14183287
- Primary Topic
- Face recognition and analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00