Enhanced human posture recognition using multi-stream deep CNN with blended ensemble classification
Abstract Human posture recognition has been an eye-catching research area for over a decade, being highly valuable for tasks such as sitting, standing, lying, and walking postures, especially in elderly care and public health monitoring. Whereas existing approaches like sensor-based, vision-based, and feature-based methods have been widely utilized, the growing availability of affordable computational power has shifted consideration toward deep learning techniques for improved accuracy and efficiency. To address these limitations, this paper presents a multi-stream deep CNN ensemble framework for reliable human posture recognition. The proposed method is novel in its design, integrating multi-stream feature aggregation from multiple pre-trained CNN backbones, PCA-based dimensionality reduction before ensemble learning, and a heterogeneous blending framework that combines complementary base classifiers. This jointly optimizes feature diversity and decision robustness, resulting competitive-to-superior accuracy on multiple benchmark posture recognition datasets relative to the pretrained CNN baselines and ensemble variants evaluated in this study. The method was evaluated against several deep learning and ensemble techniques on five publicly available datasets: KARD, MCF, NUCLA, URFD, and UP Fall. Across most evaluated datasets and metrics it matched or outperformed the compared methods in precision, accuracy, and recall, with two exceptions (NUCLA and UP Fall Front) where individual classifiers attained a marginally higher single-metric accuracy while the proposed approach achieved a more balanced precision, recall and F1-Score. The results highlight the model’s effectiveness and efficiency, highlighting its potential for real-world applications in posture recognition. Notably, the method achieves accuracies of up to 98.51%, reduces combined feature dimensionality by 93.75% through PCA, and lowers inference latency by 10.7% compared to a standalone ResNet-101 baseline measured on the same hardware (see “Environment” section and Table 12). Its improved performance as well as reduced computational requirements make it a promising candidate for sensitive use cases, such as remotely monitoring the elderly to help mitigate health risks and support access to care. Given its lower feature dimensionality and reduced inference latency relative to the standalone CNN baselines evaluated here, the framework may also be a reasonable starting point for exploring deployment on portable and edge computing devices, though this has not been directly benchmarked on such hardware.
Authors
- Khalid Munawar (ORCID: https://orcid.org/0000-0003-1557-2629)
- Syed Farooq Ali (ORCID: https://orcid.org/0000-0003-3943-903X)
- Afifa Hameed (ORCID: https://orcid.org/0000-0002-6413-5444)
- Mudassar Ali
- Muhammad Bilal
- Aaima Parvez
- Amer Hamza Aamir
- Ahmed Hasnain Mirza
Institutions
- King Abdulaziz University (SA)
- University of Central Punjab (PK)
- University of Management and Technology (PK)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1038/s41598-026-72874-4
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00