Aligning Acoustic and Semantic Representations for Speech Aspect-Based Sentiment Analysis

Fine-grained sentiment information about specific aspects is the target of aspect-based sentiment analysis (ABSA). Nevertheless, when real-world situations involve sentiment expressed through multiple modalities, traditional text-centric approaches are prone to information loss and ambiguity. In this work, we introduce a novel speech aspect-based sentiment analysis task. By taking raw speech signals as input, this task outputs aspect terms, opinion expressions and sentiment polarities associated with each aspect and thereby leverages the rich information in speech, including paralinguistic cues such as prosody, stress and intonation, to strengthen sentiment analysis. To facilitate controlled evaluation, we construct dedicated speech ABSA datasets from two domains (Restaurant and Laptop) with roughly 6 h of audio and over 9000 annotated quadruples. Text-to-speech synthesis is employed for training and validation sets in our data construction strategy, while human-recorded speech is reserved for test sets to assess generalisation under read-speech conditions. Furthermore, we adopt a cross-modal mixup approach based on optimal transport to jointly extract sentiment elements from both speech and ASR-generated text and effectively align acoustic and semantic representations. Experimental results on both domains confirm the significance of the suggested speech aspect-based sentiment analysis task, with our model reaching F1 scores of 45.28% and 40.06% on Restaurant and Laptop, respectively, and outperforming several strong audio-only, ASR-based and multimodal baselines.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-30
DOI
https://doi.org/10.3390/electronics15194483
Primary Topic
Sentiment Analysis and Opinion Mining
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Aligning Acoustic and Semantic Representations for Speech Aspect-Based Sentiment Analysis

Guodong Zhou, Zhongqing Wang, Yifei Miao
Electronics
Sentiment Analysis and Opinion Mining
article

Aligning Acoustic and Semantic Representations for Speech Aspect-Based Sentiment Analysis

Guodong Zhou, Zhongqing Wang, Yifei Miao
article en

Abstract

Fine-grained sentiment information about specific aspects is the target of aspect-based sentiment analysis (ABSA). Nevertheless, when real-world situations involve sentiment expressed through multiple modalities, traditional text-centric approaches are prone to information loss and ambiguity. In this work, we introduce a novel speech aspect-based sentiment analysis task. By taking raw speech signals as input, this task outputs aspect terms, opinion expressions and sentiment polarities associated with each aspect and thereby leverages the rich information in speech, including paralinguistic cues such as prosody, stress and intonation, to strengthen sentiment analysis. To facilitate controlled evaluation, we construct dedicated speech ABSA datasets from two domains (Restaurant and Laptop) with roughly 6 h of audio and over 9000 annotated quadruples. Text-to-speech synthesis is employed for training and validation sets in our data construction strategy, while human-recorded speech is reserved for test sets to assess generalisation under read-speech conditions. Furthermore, we adopt a cross-modal mixup approach based on optimal transport to jointly extract sentiment elements from both speech and ASR-generated text and effectively align acoustic and semantic representations. Experimental results on both domains confirm the significance of the suggested speech aspect-based sentiment analysis task, with our model reaching F1 scores of 45.28% and 40.06% on Restaurant and Laptop, respectively, and outperforming several strong audio-only, ASR-based and multimodal baselines.

ElectronicsVol. 15(19)
Soochow University (CN)
Openalex Percentile: Top 9%
Sentiment Analysis and Opinion Mining
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Aligning Acoustic and Semantic Representations for Speech Aspect-Based Sentiment Analysis — Guodong Zhou, Zhongqing Wang, et al. · Electronics (2026) | TGRS Research Map | TGRS