Confidence-Aware Semi-Supervised Vision–Language Contrastive Learning for Abnormal Behavior Recognition

Reliable abnormal behavior recognition from surveillance videos is hindered by the high cost of clip-level annotation, the scarcity of abnormal samples, and the context-dependent nature of behavioral semantics. Although vision–language models offer strong semantic transferability, their application under limited supervision remains susceptible to noisy pseudo-labels and confirmation bias. We propose confidence-aware semi-supervised vision–language contrastive learning (CA-VLC), which jointly exploits limited labeled videos and abundant unlabeled videos. Building on an existing CLIP-initialized temporal backbone, CA-VLC combines behavior-only and context-enriched text prototypes through confidence- and agreement-guided semantic fusion. For unlabeled videos, the model generates predictions from weakly augmented views and selects reliable pseudo-labels using entropy-based confidence estimation and class-adaptive thresholds. Detached weak-view targets then supervise strongly augmented views through confidence-weighted self-training without requiring an additional teacher network. Furthermore, cross-view consistency regularization and confidence-aware contextual alignment suppress unreliable semantic cues and improve robustness to contextual noise. Experiments on CABR50 demonstrate consistent improvements across multiple labeled-data ratios, while evaluations on CABRZ6 and UCF-101 assess prompt-based transfer to predefined target label sets without target-domain fine-tuning. With 10% labeled videos, CA-VLC achieves 84.06% Top-1 accuracy and 83.51% Macro-F1, retaining 95.47% of its fully supervised Top-1 accuracy of 88.05%, thereby demonstrating its effectiveness for label-efficient abnormal behavior recognition.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-09-10
DOI
https://doi.org/10.3390/info17090879
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Confidence-Aware Semi-Supervised Vision–Language Contrastive Learning for Abnormal Behavior Recognition

H.L. Liu, Jianxin Sun, Xianmin Zhao
Information
Human Pose and Action Recognition
article

Confidence-Aware Semi-Supervised Vision–Language Contrastive Learning for Abnormal Behavior Recognition

H.L. Liu, Jianxin Sun, Xianmin Zhao
article en

Abstract

Reliable abnormal behavior recognition from surveillance videos is hindered by the high cost of clip-level annotation, the scarcity of abnormal samples, and the context-dependent nature of behavioral semantics. Although vision–language models offer strong semantic transferability, their application under limited supervision remains susceptible to noisy pseudo-labels and confirmation bias. We propose confidence-aware semi-supervised vision–language contrastive learning (CA-VLC), which jointly exploits limited labeled videos and abundant unlabeled videos. Building on an existing CLIP-initialized temporal backbone, CA-VLC combines behavior-only and context-enriched text prototypes through confidence- and agreement-guided semantic fusion. For unlabeled videos, the model generates predictions from weakly augmented views and selects reliable pseudo-labels using entropy-based confidence estimation and class-adaptive thresholds. Detached weak-view targets then supervise strongly augmented views through confidence-weighted self-training without requiring an additional teacher network. Furthermore, cross-view consistency regularization and confidence-aware contextual alignment suppress unreliable semantic cues and improve robustness to contextual noise. Experiments on CABR50 demonstrate consistent improvements across multiple labeled-data ratios, while evaluations on CABRZ6 and UCF-101 assess prompt-based transfer to predefined target label sets without target-domain fine-tuning. With 10% labeled videos, CA-VLC achieves 84.06% Top-1 accuracy and 83.51% Macro-F1, retaining 95.47% of its fully supervised Top-1 accuracy of 88.05%, thereby demonstrating its effectiveness for label-efficient abnormal behavior recognition.

InformationVol. 17(9)
Chongqing Technology and Business University (CN), Chongqing University of Technology (CN)
Quality Education
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.