Design of monocular vision human arm-robotic arm interaction system based on attention feature fusion and sequential graph convolution

Abstract Aiming at the problems of insufficient spatial context modeling and limited understanding of temporal dynamics in monocular vision-based human-arm-robot interaction, this study proposes a monocular vision-based human-arm-robot interaction system based on attention feature fusion and temporal graph convolution. The system enhances spatial feature expression through the scaled dot-product self-attention mechanism guided by multi-scale feature pyramid. Moreover, it combines the video-based monocular three-dimensional pose estimation model to drive the spatial-temporal graph convolutional network to achieve intent recognition, and builds a two-level architecture of “feature enhancement-spatial and temporal evolution”. Experimental results reveals that under the optimal parameter configuration, the system achieves a precision rate of 96.7%, a recall rate of 94.1%, and an F1-Score of 95.4% in the visual perception task. Moreover, all indicators are significantly better than the comparison method ( p < 0.001). The model inference time is shortened to 20.8 ms, and the calculation amount is 58.7 GFLOPs. The trajectory tracking error is reduced to 1.18 cm, and the average multi-task grasping success rate reaches 93.2%. Meanwhile, CPU resource utilization is 58.7 GFLOPs, and energy consumption per task is controlled at 53.8 J. The proposed method effectively solves the problem of accurate perception of continuous human arm movement intentions and smooth interaction of robotic arms under monocular vision through the collaborative mechanism of multi-scale spatial feature adaptive enhancement and dynamic spatiotemporal map modeling. This provides key technical support for natural human-computer collaboration.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-03
DOI
https://doi.org/10.1038/s41598-026-73045-1
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Design of monocular vision human arm-robotic arm interaction system based on attention feature fusion and sequential graph convolution

HE Yaqin, Zhang Fei
Scientific Reports
Human Pose and Action Recognition
article

Design of monocular vision human arm-robotic arm interaction system based on attention feature fusion and sequential graph convolution

HE Yaqin, Zhang Fei
article en

Abstract

Abstract Aiming at the problems of insufficient spatial context modeling and limited understanding of temporal dynamics in monocular vision-based human-arm-robot interaction, this study proposes a monocular vision-based human-arm-robot interaction system based on attention feature fusion and temporal graph convolution. The system enhances spatial feature expression through the scaled dot-product self-attention mechanism guided by multi-scale feature pyramid. Moreover, it combines the video-based monocular three-dimensional pose estimation model to drive the spatial-temporal graph convolutional network to achieve intent recognition, and builds a two-level architecture of “feature enhancement-spatial and temporal evolution”. Experimental results reveals that under the optimal parameter configuration, the system achieves a precision rate of 96.7%, a recall rate of 94.1%, and an F1-Score of 95.4% in the visual perception task. Moreover, all indicators are significantly better than the comparison method ( p < 0.001). The model inference time is shortened to 20.8 ms, and the calculation amount is 58.7 GFLOPs. The trajectory tracking error is reduced to 1.18 cm, and the average multi-task grasping success rate reaches 93.2%. Meanwhile, CPU resource utilization is 58.7 GFLOPs, and energy consumption per task is controlled at 53.8 J. The proposed method effectively solves the problem of accurate perception of continuous human arm movement intentions and smooth interaction of robotic arms under monocular vision through the collaborative mechanism of multi-scale spatial feature adaptive enhancement and dynamic spatiotemporal map modeling. This provides key technical support for natural human-computer collaboration.

Scientific Reports
Changzhou Institute of Mechatronic Technology (CN)
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.