Asynchronous Acoustic–Image Classification for Unattended Ground Sensors with Domain Adaptation and Bidirectional Prototype Distillation

Acoustic and image sensors in unattended ground sensor systems often operate at different sampling rates and use different trigger mechanisms and fields of view, so their observations are asynchronous and lack instance-level correspondence. In addition, web images used to supplement scarce field annotations exhibit a substantial domain shift from target-domain images. This paper proposes a three-stream joint learning framework that integrates reliability-weighted image-domain adaptation, reliability-aware dynamic bidirectional prototype distillation, and balanced complementary modality dropout. A lightweight shared backbone processes labeled acoustic spectra, labeled mixed-source images, and unlabeled target-domain images. A main classifier and two auxiliary classifiers estimate target pseudo-label confidence and agreement. Their consensus is used to construct reliability weights for marginal alignment, while confidence guides target-sample selection and weighting for class-conditional alignment. Acoustic and image class prototypes transfer category-level knowledge without constructing artificial instance pairs, while direction-specific distillation weights are dynamically assigned according to classification reliability, prototype availability, and prototype stability. Across three independent runs with different random seeds, experiments on datasets comprising 148,038 acoustic spectra and 167,754 images yield acoustic, image, and overall accuracies of 91.97±0.08%, 95.16±0.14%, and 93.72±0.07%, respectively, exceeding the highest corresponding baseline mean accuracies by 1.52, 1.30, and 2.31 percentage points. The deployment model contains 123.40 K parameters and requires 0.47 MiB in FP32, making it a compact classification model with potential for deployment in resource-constrained field environments.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-25
DOI
https://doi.org/10.3390/app16199568
Primary Topic
Music and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Asynchronous Acoustic–Image Classification for Unattended Ground Sensors with Domain Adaptation and Bidirectional Prototype Distillation

Yan Wang, Bo Li, Jun Pei, Lijie Xia
Applied Sciences
Music and Audio Processing
article

Asynchronous Acoustic–Image Classification for Unattended Ground Sensors with Domain Adaptation and Bidirectional Prototype Distillation

Yan Wang, Bo Li, Jun Pei, Lijie Xia
article en

Abstract

Acoustic and image sensors in unattended ground sensor systems often operate at different sampling rates and use different trigger mechanisms and fields of view, so their observations are asynchronous and lack instance-level correspondence. In addition, web images used to supplement scarce field annotations exhibit a substantial domain shift from target-domain images. This paper proposes a three-stream joint learning framework that integrates reliability-weighted image-domain adaptation, reliability-aware dynamic bidirectional prototype distillation, and balanced complementary modality dropout. A lightweight shared backbone processes labeled acoustic spectra, labeled mixed-source images, and unlabeled target-domain images. A main classifier and two auxiliary classifiers estimate target pseudo-label confidence and agreement. Their consensus is used to construct reliability weights for marginal alignment, while confidence guides target-sample selection and weighting for class-conditional alignment. Acoustic and image class prototypes transfer category-level knowledge without constructing artificial instance pairs, while direction-specific distillation weights are dynamically assigned according to classification reliability, prototype availability, and prototype stability. Across three independent runs with different random seeds, experiments on datasets comprising 148,038 acoustic spectra and 167,754 images yield acoustic, image, and overall accuracies of 91.97±0.08%, 95.16±0.14%, and 93.72±0.07%, respectively, exceeding the highest corresponding baseline mean accuracies by 1.52, 1.30, and 2.31 percentage points. The deployment model contains 123.40 K parameters and requires 0.47 MiB in FP32, making it a compact classification model with potential for deployment in resource-constrained field environments.

Applied SciencesVol. 16(19)
Shanghai Institute of Microsystem and Information Technology (CN), University of Chinese Academy of Sciences (CN)
Openalex Percentile: Top 10%
Music and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.