Improved LinkNet with new aspect term extraction for multi-modal sarcasm detection
The rapid growth of online communication has increased the need for efficient multimodal sarcasm detection models. However, effectively integrating textual, visual, and audio information remains a significant challenge. To tackle this issue, this paper proposes a Conditional Random Field-assisted LinkNet framework (CRF-LN-MSD) that integrates aspect-level textual representations via a Modified Fisher Score-based Aspect Term Extraction (MFS-ATE) technique with facial expression features and audio spectral characteristics. The CRF-LN architecture employs a CRF layer for structured multimodal fusion and a novel Tanh-Swish-Sigmoid (TSS) activation function for improved gradient flow. Experimental results on the M2H2 and CMU-MOSI datasets demonstrate competitive performance under the evaluated experimental conditions, with a precision of 0.978 on M2H2 and a cross-validated accuracy of 0.9787 on CMU-MOSI. These results demonstrate the effectiveness of the proposed architectural refinement for multimodal sarcasm detection within the evaluated scope.
Authors
- Ramgopal Kashyap (ORCID: https://orcid.org/0000-0002-5352-1286)
- Jameer Kotwal (ORCID: https://orcid.org/0009-0000-3779-6573)
- Dnyaneshwar Bavkar
- Navnath Pokale
- Anil Pise
- Asma Shaikh
- Neha Nagdeve
Institutions
- Guru Ghasidas Vishwavidyalaya (IN)
- Centre for Artificial Intelligence and Robotics (IN)
- Mahatma Gandhi Mission's Dental College and Hospital (IN)
- Raisoni Group of Institutions (IN)
- Right to Care (ZA)
- MIT Art, Design and Technology University (IN)
Publication Details
- Journal
- Discover Applied Sciences
- Published
- 2026-09-08
- DOI
- https://doi.org/10.1007/s42452-026-09501-4
- Primary Topic
- Sentiment Analysis and Opinion Mining
- Type
- article
- Field-Weighted Citation Impact
- 0.00