Open-ended Human Activity Understanding via LLM-assisted Motion Decomposition and Semantic Fusion

Human activity is a fundamental component of intelligent interactive systems. Existing human activity recognition (HAR) follows a closed-ended assumption—defining a fixed set of activity classes and classifying each sensor signal into one of them—yet real-world human behavior is complex and ever-changing and thus cannot be fully represented by any fixed set. We present a new paradigm for open-ended human activity understanding (HAU), shifting the goal from closed-set classification to understanding activities via their meta-motions. To achieve this, we design MOSAIC, a motion decomposition and semantic fusion framework that transforms raw sensor streams into structured meta-motion representations and subsequently generates fine-grained natural-language descriptions. Additionally, we propose an LLM-driven training strategy for sensor-language alignment, enabling effective cross-modal supervision in the absence of real-world training data. Finally, we leverage LLM-assisted reasoning to recognize complex high-level activities based on the understanding of underlying motions and their composition. Our model demonstrates strong generalization and performance across 18 public HAR datasets, outperforming the best baseline by up to 19.6% in unseen scenarios.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Published
2026-09-30
DOI
https://doi.org/10.1145/3831643
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Open-ended Human Activity Understanding via LLM-assisted Motion Decomposition and Semantic Fusion

Jiaming Huang, Cheng Chao Guo, Wei Dong, Qingxin Wei et al.
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Human Pose and Action Recognition
article

Open-ended Human Activity Understanding via LLM-assisted Motion Decomposition and Semantic Fusion

Jiaming Huang, Cheng Chao Guo, Wei Dong, Qingxin Wei, Gao Y, Kai Hu
article en

Abstract

Human activity is a fundamental component of intelligent interactive systems. Existing human activity recognition (HAR) follows a closed-ended assumption—defining a fixed set of activity classes and classifying each sensor signal into one of them—yet real-world human behavior is complex and ever-changing and thus cannot be fully represented by any fixed set. We present a new paradigm for open-ended human activity understanding (HAU), shifting the goal from closed-set classification to understanding activities via their meta-motions. To achieve this, we design MOSAIC, a motion decomposition and semantic fusion framework that transforms raw sensor streams into structured meta-motion representations and subsequently generates fine-grained natural-language descriptions. Additionally, we propose an LLM-driven training strategy for sensor-language alignment, enabling effective cross-modal supervision in the absence of real-world training data. Finally, we leverage LLM-assisted reasoning to recognize complex high-level activities based on the understanding of underlying motions and their composition. Our model demonstrates strong generalization and performance across 18 public HAR datasets, outperforming the best baseline by up to 19.6% in unseen scenarios.

Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous TechnologiesVol. 10(3)
Zhejiang University (CN)
Openalex Percentile: Top 15%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.