ActNet: focus-aware multi-scale CNN for human activity recognition from images
Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address this problem. ActNet integrates a Multi-Feature Network (MFNet) backbone for extracting rich features from multiple receptive fields, an Activity Multi-scale Block (AMB) for learning spatially diverse action patterns, and a Focus-Aware Recognition Module (FARM) that adaptively highlights the most informative regions of the image. We evaluate ActNet on the Stanford 40 Actions and PASCAL Visual Object Classes (VOC) 2012 datasets and show that it outperforms several state-of-the-art CNN and transformer-based models, achieving superior accuracy, precision, recall, and F1-score. Extensive ablation studies confirm the effectiveness of both AMB and FARM components. ActNet demonstrates robust generalization to a wide range of human actions, making it a strong candidate for still-image-based action recognition tasks in practical applications.
Authors
- Şafak Kılıç (ORCID: https://orcid.org/0000-0002-2014-7638)
Institutions
- University of Nottingham (GB)
- Kayseri Eğitim ve Araştırma Hastanesi (TR)
Publication Details
- Journal
- PeerJ Computer Science
- Published
- 2026-09-14
- DOI
- https://doi.org/10.7717/peerj-cs.3990
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00