Domain adaptation and self-training for improved cross-batch classification of spectroscopic data: a comparative analysis of feature normalization techniques

Surface-Enhanced Raman Spectroscopy (SERS) enables sensitive, label-free chemical identification, but machine-learning models trained on a single acquisition session frequently underperform when applied to spectra from a different batch, owing to substrate, instrument, and session variability. We study this cross-session domain shift problem on the publicly available Rhodamine 6G (R6G) SERS dataset released by Park et al. [1], which contains two batches (Batch-1 and Batch-2; 1500 spectra each) collected under nominally similar conditions but with clear differences in intensity scale, baseline behavior, and concentration range. We propose a domain-adaptive feature-learning pipeline that combines a transformer-based denoising autoencoder with Correlation Alignment in the latent space (T-DAE+CORAL) to reduce inter-batch discrepancies without requiring target labels. In the more challenging transfer direction (B1→B2, where the no-adaptation baseline is lowest), a logistic regression classifier on T-DAE+CORAL latent features improves balanced accuracy from 0.844 to 0.884 ± 0.019, with specificity rising from 0.688 to 0.939. In the reverse direction (B2→B1), adding 100 labeled target spectra (~ 6.7% of the target batch) in a semi-supervised setting raises balanced accuracy from 0.887 to 0.955 ± 0.044. We frame this work as a methodological proof-of-concept for cross-batch alignment in SERS classification using a single benchmark analyte; broader validation across additional analytes, biomolecular targets, and multi-instrument datasets is identified as a priority direction for future work.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-09
DOI
https://doi.org/10.1038/s41598-026-61856-1
Primary Topic
Spectroscopy Techniques in Biomedical and Chemical Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Domain adaptation and self-training for improved cross-batch classification of spectroscopic data: a comparative analysis of feature normalization techniques

Molla Ehsanul Majid, Muhammad EH. Chowdhury, Mehmet Burçin Ünlü, Md. Sakib Bin Islam et al.
Scientific Reports
Spectroscopy Techniques in Biomedical and Chemical Research
article

Domain adaptation and self-training for improved cross-batch classification of spectroscopic data: a comparative analysis of feature normalization techniques

Molla Ehsanul Majid, Muhammad EH. Chowdhury, Mehmet Burçin Ünlü, Md. Sakib Bin Islam, Gozde Durmus, Saad B. A. Kashem, Amith Khandakar
article en

Abstract

Surface-Enhanced Raman Spectroscopy (SERS) enables sensitive, label-free chemical identification, but machine-learning models trained on a single acquisition session frequently underperform when applied to spectra from a different batch, owing to substrate, instrument, and session variability. We study this cross-session domain shift problem on the publicly available Rhodamine 6G (R6G) SERS dataset released by Park et al. [1], which contains two batches (Batch-1 and Batch-2; 1500 spectra each) collected under nominally similar conditions but with clear differences in intensity scale, baseline behavior, and concentration range. We propose a domain-adaptive feature-learning pipeline that combines a transformer-based denoising autoencoder with Correlation Alignment in the latent space (T-DAE+CORAL) to reduce inter-batch discrepancies without requiring target labels. In the more challenging transfer direction (B1→B2, where the no-adaptation baseline is lowest), a logistic regression classifier on T-DAE+CORAL latent features improves balanced accuracy from 0.844 to 0.884 ± 0.019, with specificity rising from 0.688 to 0.939. In the reverse direction (B2→B1), adding 100 labeled target spectra (~ 6.7% of the target batch) in a semi-supervised setting raises balanced accuracy from 0.887 to 0.955 ± 0.044. We frame this work as a methodological proof-of-concept for cross-batch alignment in SERS classification using a single benchmark analyte; broader validation across additional analytes, biomolecular targets, and multi-instrument datasets is identified as a priority direction for future work.

Scientific Reports
Academic Bridge Program (QA), Özyeğin University (TR), Qatar University (QA), Qatar Foundation (QA), Stanford University (US)
Life below water
Openalex Percentile: Top 12%
Spectroscopy Techniques in Biomedical and Chemical Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.