Improved LinkNet with new aspect term extraction for multi-modal sarcasm detection

The rapid growth of online communication has increased the need for efficient multimodal sarcasm detection models. However, effectively integrating textual, visual, and audio information remains a significant challenge. To tackle this issue, this paper proposes a Conditional Random Field-assisted LinkNet framework (CRF-LN-MSD) that integrates aspect-level textual representations via a Modified Fisher Score-based Aspect Term Extraction (MFS-ATE) technique with facial expression features and audio spectral characteristics. The CRF-LN architecture employs a CRF layer for structured multimodal fusion and a novel Tanh-Swish-Sigmoid (TSS) activation function for improved gradient flow. Experimental results on the M2H2 and CMU-MOSI datasets demonstrate competitive performance under the evaluated experimental conditions, with a precision of 0.978 on M2H2 and a cross-validated accuracy of 0.9787 on CMU-MOSI. These results demonstrate the effectiveness of the proposed architectural refinement for multimodal sarcasm detection within the evaluated scope.

Authors

Institutions

Publication Details

Journal
Discover Applied Sciences
Published
2026-09-08
DOI
https://doi.org/10.1007/s42452-026-09501-4
Primary Topic
Sentiment Analysis and Opinion Mining
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Improved LinkNet with new aspect term extraction for multi-modal sarcasm detection

Ramgopal Kashyap, Jameer Kotwal, Dnyaneshwar Bavkar, Navnath Pokale et al.
Discover Applied Sciences
Sentiment Analysis and Opinion Mining
article

Improved LinkNet with new aspect term extraction for multi-modal sarcasm detection

Ramgopal Kashyap, Jameer Kotwal, Dnyaneshwar Bavkar, Navnath Pokale, Anil Pise, Asma Shaikh, Neha Nagdeve
article en

Abstract

The rapid growth of online communication has increased the need for efficient multimodal sarcasm detection models. However, effectively integrating textual, visual, and audio information remains a significant challenge. To tackle this issue, this paper proposes a Conditional Random Field-assisted LinkNet framework (CRF-LN-MSD) that integrates aspect-level textual representations via a Modified Fisher Score-based Aspect Term Extraction (MFS-ATE) technique with facial expression features and audio spectral characteristics. The CRF-LN architecture employs a CRF layer for structured multimodal fusion and a novel Tanh-Swish-Sigmoid (TSS) activation function for improved gradient flow. Experimental results on the M2H2 and CMU-MOSI datasets demonstrate competitive performance under the evaluated experimental conditions, with a precision of 0.978 on M2H2 and a cross-validated accuracy of 0.9787 on CMU-MOSI. These results demonstrate the effectiveness of the proposed architectural refinement for multimodal sarcasm detection within the evaluated scope.

Discover Applied Sciences
Guru Ghasidas Vishwavidyalaya (IN), Centre for Artificial Intelligence and Robotics (IN), Mahatma Gandhi Mission's Dental College and Hospital (IN), Raisoni Group of Institutions (IN), Right to Care (ZA), MIT Art, Design and Technology University (IN)
Openalex Percentile: Top 8%
Sentiment Analysis and Opinion Mining
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Improved LinkNet with new aspect term extraction for multi-modal sarcasm detection — Ramgopal Kashyap, Jameer Kotwal, et al. · Discover Applied Sciences (2026) | TGRS Research Map | TGRS