Feasibility of Using a Multi-Agent LLM System to Correct Annotations and Support Low-Effort Activity Labeling

Accurate measurement of behaviors is critical for research in human-computer interaction, ubiquitous computing, and personal health informatics, because it underpins many health tracking and intervention systems. Most data collection studies, however, still rely on participants' self-reports and manual annotations, which often under- or over-estimate activity duration and type, and require substantial effort from researchers to clean and validate. An automated system that can combine passive sensing data with participants' self-reports might detect inconsistencies and suggest corrections. We introduce GLOSS4HAR, a multi-agent LLM-based system designed to mimic human sensemaking and assist researchers in cleaning and refining activity annotations. We demonstrate the potential of GLOSS4HAR in two key tasks: (1) correcting and reconciling participant self-annotations, and (2) triangulating passive sensing data with different forms of lightweight self-reports to generate accurate activity timelines. Our evaluation shows that GLOSS4HAR improves annotation quality by up to 9.9% in F1 score and can reconstruct activity timelines that align with human annotations at 75-92% F1. Based on our findings, we discuss the implications of our work for the next generation of activity annotation systems that might use human-AI collaboration.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Published
2026-09-30
DOI
https://doi.org/10.1145/3831653
Citations
1
Primary Topic
Personal Information Management and User Behavior
Type
article
Field-Weighted Citation Impact
10.24
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Feasibility of Using a Multi-Agent LLM System to Correct Annotations and Support Low-Effort Activity Labeling

Akshat Choube, Varun Mishra, STEPHEN S. INTILLE, Ha Le
1 citations
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Personal Information Management and User Behavior
10.24
article

Feasibility of Using a Multi-Agent LLM System to Correct Annotations and Support Low-Effort Activity Labeling

Akshat Choube, Varun Mishra, STEPHEN S. INTILLE, Ha Le
article en
1 citations

Abstract

Accurate measurement of behaviors is critical for research in human-computer interaction, ubiquitous computing, and personal health informatics, because it underpins many health tracking and intervention systems. Most data collection studies, however, still rely on participants' self-reports and manual annotations, which often under- or over-estimate activity duration and type, and require substantial effort from researchers to clean and validate. An automated system that can combine passive sensing data with participants' self-reports might detect inconsistencies and suggest corrections. We introduce GLOSS4HAR, a multi-agent LLM-based system designed to mimic human sensemaking and assist researchers in cleaning and refining activity annotations. We demonstrate the potential of GLOSS4HAR in two key tasks: (1) correcting and reconciling participant self-annotations, and (2) triangulating passive sensing data with different forms of lightweight self-reports to generate accurate activity timelines. Our evaluation shows that GLOSS4HAR improves annotation quality by up to 9.9% in F1 score and can reconstruct activity timelines that align with human annotations at 75-92% F1. Based on our findings, we discuss the implications of our work for the next generation of activity annotation systems that might use human-AI collaboration.

Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous TechnologiesVol. 10(3)
Northeastern University (US)
Industry, innovation and infrastructure
Openalex Percentile: Top 1%
Personal Information Management and User Behavior
10.24
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.