Monocular Interacting-Hand Reconstruction with Multimodal Context Fusion and Spatial Attention

3D hand reconstruction from monocular RGB images has attracted increasing attention due to its low cost and ease of deployment. However, accurately reconstructing interacting hands remains challenging, primarily because mutual occlusion leads to missing visual evidence, while image cropping weakens global spatial consistency. To address these issues, we propose the Context-Aware Interacting Hand Reconstruction Network (CANet), a monocular interacting-hand reconstruction framework that integrates multimodal context fusion with structured spatial attention. Specifically, CANet leverages estimated depth maps, edge contours, and center heatmaps to retain global contextual cues and guide reconstruction in occluded regions. A hierarchical spatial attention module further enhances spatial consistency by separately modeling intra-hand structural dependencies and inter-hand interactions, enabling more coherent reasoning about complex hand poses. Experiments across multiple public benchmarks demonstrate that CANet consistently improves the accuracy of both hand pose estimation and mesh reconstruction.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Multimedia Computing Communications and Applications
Published
2026-09-30
DOI
https://doi.org/10.1145/3845997
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Monocular Interacting-Hand Reconstruction with Multimodal Context Fusion and Spatial Attention

Yuchun Fang, Hanzhang Wang, Yiting Cao, Junfeng Zhu et al.
ACM Transactions on Multimedia Computing Communications and Applications
Human Pose and Action Recognition
article

Monocular Interacting-Hand Reconstruction with Multimodal Context Fusion and Spatial Attention

Yuchun Fang, Hanzhang Wang, Yiting Cao, Junfeng Zhu, Cheng Jin, Te Li
article en

Abstract

3D hand reconstruction from monocular RGB images has attracted increasing attention due to its low cost and ease of deployment. However, accurately reconstructing interacting hands remains challenging, primarily because mutual occlusion leads to missing visual evidence, while image cropping weakens global spatial consistency. To address these issues, we propose the Context-Aware Interacting Hand Reconstruction Network (CANet), a monocular interacting-hand reconstruction framework that integrates multimodal context fusion with structured spatial attention. Specifically, CANet leverages estimated depth maps, edge contours, and center heatmaps to retain global contextual cues and guide reconstruction in occluded regions. A hierarchical spatial attention module further enhances spatial consistency by separately modeling intra-hand structural dependencies and inter-hand interactions, enabling more coherent reasoning about complex hand poses. Experiments across multiple public benchmarks demonstrate that CANet consistently improves the accuracy of both hand pose estimation and mesh reconstruction.

ACM Transactions on Multimedia Computing Communications and Applications
Shanghai University (CN)
Sustainable cities and communities
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Monocular Interacting-Hand Reconstruction with Multimodal Context Fusion and Spatial Attention — Yuchun Fang, Hanzhang Wang, et al. · ACM Transactions on Multimedia Computing Communications and Applications (2026) | TGRS Research Map | TGRS