MG2CL: Multi-Granularity Graph Contrastive Learning for Skeleton-Based Action Recognition

Skeleton-based action recognition has gained increasing attention as a robust alternative to RGB-based methods due to its resilience to background clutter and lighting variations. Existing approaches have achieved notable progress, particularly with graph-based models that capture topological and temporal patterns of human joints. However, most methods still struggle to model relationships between non-naturally connected joints and fail to fully exploit the fine-grained spatio-temporal characteristics of skeleton sequences, limiting generalization and recognition accuracy. To address these limitations, we propose MG2CL, a Multi-Granularity Graph Contrastive Learning framework designed to enhance skeleton representation learning. Our approach employs a multi-granularity graph structure to model both short- and long-range joint interactions through complementary body-part relations. Furthermore, we introduce a novel spatio-temporal masking strategy for data augmentation, encouraging the model to learn more diverse and informative patterns. A semantic-level memory bank, built upon the multi-granularity graph, is integrated to reinforce the model’s ability to distinguish subtle action variations. Extensive experiments show that MG2CL achieves competitive performance on widely used benchmark datasets, including NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA. These results demonstrate the framework’s strong generalization and discriminative capability. Our findings suggest that incorporating multi-granularity structural information and contrastive learning principles can lead to more robust and flexible skeleton-based action models. MG2CL offers a promising direction for future work in representation learning for spatio-temporal graph data, with potential applications in human-computer interaction, surveillance, and healthcare.

Authors

Institutions

Publication Details

Journal
Transactions on Graph Intelligence and Network Applications
Published
2026-09-22
DOI
https://doi.org/10.53941/tgina.2026.100007
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MG2CL: Multi-Granularity Graph Contrastive Learning for Skeleton-Based Action Recognition

Xin Chen, Thomas Weise, Zhize Wu, Shengwei Ji et al.
Transactions on Graph Intelligence and Network Applications
Human Pose and Action Recognition
article

MG2CL: Multi-Granularity Graph Contrastive Learning for Skeleton-Based Action Recognition

Xin Chen, Thomas Weise, Zhize Wu, Shengwei Ji, Pensong Wang, Fei Liu
article en

Abstract

Skeleton-based action recognition has gained increasing attention as a robust alternative to RGB-based methods due to its resilience to background clutter and lighting variations. Existing approaches have achieved notable progress, particularly with graph-based models that capture topological and temporal patterns of human joints. However, most methods still struggle to model relationships between non-naturally connected joints and fail to fully exploit the fine-grained spatio-temporal characteristics of skeleton sequences, limiting generalization and recognition accuracy. To address these limitations, we propose MG2CL, a Multi-Granularity Graph Contrastive Learning framework designed to enhance skeleton representation learning. Our approach employs a multi-granularity graph structure to model both short- and long-range joint interactions through complementary body-part relations. Furthermore, we introduce a novel spatio-temporal masking strategy for data augmentation, encouraging the model to learn more diverse and informative patterns. A semantic-level memory bank, built upon the multi-granularity graph, is integrated to reinforce the model’s ability to distinguish subtle action variations. Extensive experiments show that MG2CL achieves competitive performance on widely used benchmark datasets, including NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA. These results demonstrate the framework’s strong generalization and discriminative capability. Our findings suggest that incorporating multi-granularity structural information and contrastive learning principles can lead to more robust and flexible skeleton-based action models. MG2CL offers a promising direction for future work in representation learning for spatio-temporal graph data, with potential applications in human-computer interaction, surveillance, and healthcare.

Transactions on Graph Intelligence and Network Applications
Anhui University (CN), Hefei University of Technology (CN), Hefei University (CN)
Reduced inequalities
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.