MFA-Pose: Human Pose Estimation with Multi-Scale Context Fusion and Adaptive Gated Upsampling for Industrial Surveillance

Human pose estimation in factory surveillance is challenged by scale variation, limb occlusion, complex backgrounds, and spatial detail loss during upsampling. This paper proposes MFA-Pose, an improved YOLO11s-Pose framework that enhances contextual representation and cross-scale feature reconstruction. The Multi-Scale Context Fusion (MSCF) module preserves directional positional information and models multi-range cross-channel dependencies using parallel one-dimensional convolutions with kernel sizes of 3, 5, and 7. The complete C2MSCF module contains approximately 0.790 M learnable parameters, compared with approximately 0.991 M parameters in the original C2PSA module. The Multi-Receptive-Field Adaptive Gated Upsampling (MRAG) module predicts sampling offsets through standard and dilated convolution branches, suppresses unreliable offsets using gating mechanisms, and combines dynamic resampling with a stable bilinear interpolation reference. The two MRAG modules used in MFA-Pose contain approximately 0.457 M and 0.130 M learnable parameters, respectively, corresponding to approximately 0.587 M parameters in total. On COCO 2017, MFA-Pose achieves 89.0% AP50pose and 61.7% AP50:95pose, outperforming YOLO11s-Pose by 2.7 and 1.6 percentage points, respectively, with only 0.40 M additional parameters. Quantitative evaluation on the annotated industrial surveillance test set further shows that MFA-Pose achieves 87.2% Precision, 77.0% Recall, 81.8% F1-score, 88.7% AP50pose, and 61.1% AP50:95pose, outperforming the YOLO11s-Pose baseline across all evaluated accuracy metrics. Qualitative comparisons further show more coherent pose predictions under self-occlusion, non-standard working postures, and cluttered equipment backgrounds. Overall, MFA-Pose provides a favorable balance between pose estimation accuracy and model complexity for industrial surveillance applications.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-04
DOI
https://doi.org/10.3390/electronics15173993
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MFA-Pose: Human Pose Estimation with Multi-Scale Context Fusion and Adaptive Gated Upsampling for Industrial Surveillance

Hanbo Zhang, Jing Huang
Electronics
Human Pose and Action Recognition
article

MFA-Pose: Human Pose Estimation with Multi-Scale Context Fusion and Adaptive Gated Upsampling for Industrial Surveillance

Hanbo Zhang, Jing Huang
article en

Abstract

Human pose estimation in factory surveillance is challenged by scale variation, limb occlusion, complex backgrounds, and spatial detail loss during upsampling. This paper proposes MFA-Pose, an improved YOLO11s-Pose framework that enhances contextual representation and cross-scale feature reconstruction. The Multi-Scale Context Fusion (MSCF) module preserves directional positional information and models multi-range cross-channel dependencies using parallel one-dimensional convolutions with kernel sizes of 3, 5, and 7. The complete C2MSCF module contains approximately 0.790 M learnable parameters, compared with approximately 0.991 M parameters in the original C2PSA module. The Multi-Receptive-Field Adaptive Gated Upsampling (MRAG) module predicts sampling offsets through standard and dilated convolution branches, suppresses unreliable offsets using gating mechanisms, and combines dynamic resampling with a stable bilinear interpolation reference. The two MRAG modules used in MFA-Pose contain approximately 0.457 M and 0.130 M learnable parameters, respectively, corresponding to approximately 0.587 M parameters in total. On COCO 2017, MFA-Pose achieves 89.0% AP50pose and 61.7% AP50:95pose, outperforming YOLO11s-Pose by 2.7 and 1.6 percentage points, respectively, with only 0.40 M additional parameters. Quantitative evaluation on the annotated industrial surveillance test set further shows that MFA-Pose achieves 87.2% Precision, 77.0% Recall, 81.8% F1-score, 88.7% AP50pose, and 61.1% AP50:95pose, outperforming the YOLO11s-Pose baseline across all evaluated accuracy metrics. Qualitative comparisons further show more coherent pose predictions under self-occlusion, non-standard working postures, and cluttered equipment backgrounds. Overall, MFA-Pose provides a favorable balance between pose estimation accuracy and model complexity for industrial surveillance applications.

ElectronicsVol. 15(17)
Zhejiang Sci-Tech University (CN)
Industry, innovation and infrastructure
Openalex Percentile: Top 12%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.