ActNet: focus-aware multi-scale CNN for human activity recognition from images

Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address this problem. ActNet integrates a Multi-Feature Network (MFNet) backbone for extracting rich features from multiple receptive fields, an Activity Multi-scale Block (AMB) for learning spatially diverse action patterns, and a Focus-Aware Recognition Module (FARM) that adaptively highlights the most informative regions of the image. We evaluate ActNet on the Stanford 40 Actions and PASCAL Visual Object Classes (VOC) 2012 datasets and show that it outperforms several state-of-the-art CNN and transformer-based models, achieving superior accuracy, precision, recall, and F1-score. Extensive ablation studies confirm the effectiveness of both AMB and FARM components. ActNet demonstrates robust generalization to a wide range of human actions, making it a strong candidate for still-image-based action recognition tasks in practical applications.

Authors

Institutions

Publication Details

Journal
PeerJ Computer Science
Published
2026-09-14
DOI
https://doi.org/10.7717/peerj-cs.3990
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ActNet: focus-aware multi-scale CNN for human activity recognition from images

Şafak Kılıç
PeerJ Computer Science
Human Pose and Action Recognition
article

ActNet: focus-aware multi-scale CNN for human activity recognition from images

Şafak Kılıç
article en

Abstract

Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address this problem. ActNet integrates a Multi-Feature Network (MFNet) backbone for extracting rich features from multiple receptive fields, an Activity Multi-scale Block (AMB) for learning spatially diverse action patterns, and a Focus-Aware Recognition Module (FARM) that adaptively highlights the most informative regions of the image. We evaluate ActNet on the Stanford 40 Actions and PASCAL Visual Object Classes (VOC) 2012 datasets and show that it outperforms several state-of-the-art CNN and transformer-based models, achieving superior accuracy, precision, recall, and F1-score. Extensive ablation studies confirm the effectiveness of both AMB and FARM components. ActNet demonstrates robust generalization to a wide range of human actions, making it a strong candidate for still-image-based action recognition tasks in practical applications.

PeerJ Computer ScienceVol. 12
University of Nottingham (GB), Kayseri Eğitim ve Araştırma Hastanesi (TR)
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

ActNet: focus-aware multi-scale CNN for human activity recognition from images — Şafak Kılıç · PeerJ Computer Science (2026) | TGRS Research Map | TGRS