Adapting Pretrained Large Vision Models for Sensor-based Activity Recognition

Understanding and recognizing human activities from low-cost wearable sensors has attracted increasing attention in recent years. To achieve this goal, numerous learning models and augmentation approaches have been developed. While effective in certain scenarios, their performance is often limited due to the distribution shift between the training and testing data collected from different users or different body placements. Moreover, collecting large-scale and diverse sensor training data is costly and labor-intensive, which further exacerbates the data scarcity problem in HAR. To mitigate these challenges, in this work, we propose VisionHAR, a novel HAR framework that borrows knowledge from other data-rich modalities, i.e., low-cost and internet-scale image data. We first “draw” continuous sensor data on a figure to preserve both their temporal and periodic patterns. Then, we design a parameter-efficient transfer learning method to utilize the generalization capability of Large Vision Models (LVMs) pretrained on large-scale image data. To enable real-time activity recognition on edge devices, we further design a novel distillation approach to learn a highly effective and efficient student model, which achieves comparable performance with significantly fewer parameters. We evaluate our model in two main settings, i.e., cross-domain HAR and complex HAR (more than 18 activity categories). Experimental results show that VisionHAR outperforms the best existing HAR models by at least 8.22% in average accuracy and 11.48% in F1-score for cross-domain HAR with 1,460 times fewer parameters , and 7.62% in average accuracy and 7.86% in F1-score for complex HAR. Our findings suggest that the knowledge learned from vision data is generalizable to continuous sensor data, which provides a new potential solution for the data shortage issue in the activity recognition community. We release our code at https://github.com/saiketa/VisionHAR for future studies in the ubiquitous computing community.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Published
2026-09-30
DOI
https://doi.org/10.1145/3832003
Primary Topic
Context-Aware Activity Recognition Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adapting Pretrained Large Vision Models for Sensor-based Activity Recognition

Zhiqing Hong, Yunhuai Liu, Baoshen Guo, Kunlin Cai et al.
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Context-Aware Activity Recognition Systems
article

Adapting Pretrained Large Vision Models for Sensor-based Activity Recognition

Zhiqing Hong, Yunhuai Liu, Baoshen Guo, Kunlin Cai, Yize Cai, Rui Feng
article en

Abstract

Understanding and recognizing human activities from low-cost wearable sensors has attracted increasing attention in recent years. To achieve this goal, numerous learning models and augmentation approaches have been developed. While effective in certain scenarios, their performance is often limited due to the distribution shift between the training and testing data collected from different users or different body placements. Moreover, collecting large-scale and diverse sensor training data is costly and labor-intensive, which further exacerbates the data scarcity problem in HAR. To mitigate these challenges, in this work, we propose VisionHAR, a novel HAR framework that borrows knowledge from other data-rich modalities, i.e., low-cost and internet-scale image data. We first “draw” continuous sensor data on a figure to preserve both their temporal and periodic patterns. Then, we design a parameter-efficient transfer learning method to utilize the generalization capability of Large Vision Models (LVMs) pretrained on large-scale image data. To enable real-time activity recognition on edge devices, we further design a novel distillation approach to learn a highly effective and efficient student model, which achieves comparable performance with significantly fewer parameters. We evaluate our model in two main settings, i.e., cross-domain HAR and complex HAR (more than 18 activity categories). Experimental results show that VisionHAR outperforms the best existing HAR models by at least 8.22% in average accuracy and 11.48% in F1-score for cross-domain HAR with 1,460 times fewer parameters , and 7.62% in average accuracy and 7.86% in F1-score for complex HAR. Our findings suggest that the knowledge learned from vision data is generalizable to continuous sensor data, which provides a new potential solution for the data shortage issue in the activity recognition community. We release our code at https://github.com/saiketa/VisionHAR for future studies in the ubiquitous computing community.

Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous TechnologiesVol. 10(3)
University of California, Los Angeles (US), Peking University (CN), Singapore-MIT Alliance for Research and Technology (SG), The Hong Kong University of Science and Technology (Guangzhou) (CN)
Decent work and economic growth
Openalex Percentile: Top 15%
Context-Aware Activity Recognition Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.