Monocular Interacting-Hand Reconstruction with Multimodal Context Fusion and Spatial Attention
3D hand reconstruction from monocular RGB images has attracted increasing attention due to its low cost and ease of deployment. However, accurately reconstructing interacting hands remains challenging, primarily because mutual occlusion leads to missing visual evidence, while image cropping weakens global spatial consistency. To address these issues, we propose the Context-Aware Interacting Hand Reconstruction Network (CANet), a monocular interacting-hand reconstruction framework that integrates multimodal context fusion with structured spatial attention. Specifically, CANet leverages estimated depth maps, edge contours, and center heatmaps to retain global contextual cues and guide reconstruction in occluded regions. A hierarchical spatial attention module further enhances spatial consistency by separately modeling intra-hand structural dependencies and inter-hand interactions, enabling more coherent reasoning about complex hand poses. Experiments across multiple public benchmarks demonstrate that CANet consistently improves the accuracy of both hand pose estimation and mesh reconstruction.
Authors
- Yuchun Fang (ORCID: https://orcid.org/0000-0002-7085-8876)
- Hanzhang Wang (ORCID: https://orcid.org/0009-0006-4088-5984)
- Yiting Cao (ORCID: https://orcid.org/0000-0002-0990-5602)
- Junfeng Zhu (ORCID: https://orcid.org/0009-0001-7042-0727)
- Cheng Jin (ORCID: https://orcid.org/0009-0004-0063-0368)
- Te Li (ORCID: https://orcid.org/0009-0003-6059-4728)
Institutions
- Shanghai University (CN)
Publication Details
- Journal
- ACM Transactions on Multimedia Computing Communications and Applications
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1145/3845997
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00