Prompt-Controlled Multi-Modal Tuning for Person-Specific Few-Shot Micro-Expression Recognition

Micro-expression recognition (MER) is a challenging step in various multimedia applications, such as media understanding and human–computer interaction, as it can reveal genuine human emotions. However, traditional MER often overlooks person-specific facial nuances, limiting generalization and personalized adaptability in practical applications. In this paper, we propose a novel person-specific few-shot MER benchmark that uses only an image/video of the target person as a reference to extend the boundaries of generalization and personalized adaptability in MER. A Prompt-Controlled Multi-Modal Tuning (PCMMT) framework is presented to tackle this benchmark. We introduce a prompt bottleneck mechanism in PCMMT that leverages text prompts and the CLIP multi-modal embedding space to bridge the query and reference visual inputs. Our prompt bottleneck also controls the flow of information across modalities to extract person-specific, subtle micro-expression motions. Moreover, adapter groups are designed at the layer level, with multi-modal tokens fine-tuned separately to refine person-specific cues and enhance generalization. The experimental results show that the proposed person-specific few-shot setting achieves better personalized generalization than traditional MER settings, and our PCMMT outperforms previous state-of-the-art models across various MER benchmarks.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-25
DOI
https://doi.org/10.3390/s26196083
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Prompt-Controlled Multi-Modal Tuning for Person-Specific Few-Shot Micro-Expression Recognition

Jiateng Liu, Hengcan Shi, Ruiqi Wang, Kerong Li et al.
Sensors
Emotion and Mood Recognition
article

Prompt-Controlled Multi-Modal Tuning for Person-Specific Few-Shot Micro-Expression Recognition

Jiateng Liu, Hengcan Shi, Ruiqi Wang, Kerong Li, Yaonan Wang, Yingtian Yu, Tianxiang Cao
article en

Abstract

Micro-expression recognition (MER) is a challenging step in various multimedia applications, such as media understanding and human–computer interaction, as it can reveal genuine human emotions. However, traditional MER often overlooks person-specific facial nuances, limiting generalization and personalized adaptability in practical applications. In this paper, we propose a novel person-specific few-shot MER benchmark that uses only an image/video of the target person as a reference to extend the boundaries of generalization and personalized adaptability in MER. A Prompt-Controlled Multi-Modal Tuning (PCMMT) framework is presented to tackle this benchmark. We introduce a prompt bottleneck mechanism in PCMMT that leverages text prompts and the CLIP multi-modal embedding space to bridge the query and reference visual inputs. Our prompt bottleneck also controls the flow of information across modalities to extract person-specific, subtle micro-expression motions. Moreover, adapter groups are designed at the layer level, with multi-modal tokens fine-tuned separately to refine person-specific cues and enhance generalization. The experimental results show that the proposed person-specific few-shot setting achieves better personalized generalization than traditional MER settings, and our PCMMT outperforms previous state-of-the-art models across various MER benchmarks.

SensorsVol. 26(19)
Hunan University (CN), China Mobile (China) (CN), Beijing Academy of Artificial Intelligence (CN), Ministry of Education (CL), Centre for Artificial Intelligence and Robotics (IN), Southeast University (CN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.