Mixture-of-prompt experts for federated sentiment classification under non-IID data

Federated learning (FL) enables collaborative text classification while preserving data privacy, yet its performance can be degraded under non-IID data. This challenge is particularly severe in natural language processing, where client data may differ in linguistic style, topic distribution, and label semantics. Recent federated prompt-based methods address this issue by adapting large pre-trained language models (PLMs) through learnable prompts. However, most existing approaches rely on a single prompt generator, which limits their ability to capture diverse client characteristics. In this paper, we propose a novel mixture-of-prompt experts framework for federated text classification in non-IID settings. Our framework replaces the single dynamic prompt generator with multiple prompt experts that produce input-conditional prompts, and employs a gating network to dynamically combine expert outputs for each input instance. To balance generalization and personalization, the prompt experts are globally shared and aggregated via federated averaging, while the gating network is maintained locally on each client to enable client-specific expert selection without additional communication cost. This design allows the model to adapt at both the sample and client levels. We evaluate the proposed framework on the Stanford Sentiment Treebank, IMDB, and Banking77 benchmarks, spanning both binary sentiment classification and multi-class intent detection. Experimental results demonstrate that our model outperforms other baselines. Moreover, to comprehensively evaluate the proposed framework, we conduct sensitivity experiments across various hyperparameter configurations. These analyses demonstrate the stability of our model and provide insights into the trade-offs between performance and computational efficiency.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-25
DOI
https://doi.org/10.1038/s41598-026-70133-0
Primary Topic
Privacy-Preserving Technologies in Data
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Mixture-of-prompt experts for federated sentiment classification under non-IID data

Cheng-Ying Hsieh, Ching-Sheng Lin, Xin Wang, Joen-Rong Sheu et al.
Scientific Reports
Privacy-Preserving Technologies in Data
article

Mixture-of-prompt experts for federated sentiment classification under non-IID data

Cheng-Ying Hsieh, Ching-Sheng Lin, Xin Wang, Joen-Rong Sheu, Chun-Ming Yang
article en

Abstract

Federated learning (FL) enables collaborative text classification while preserving data privacy, yet its performance can be degraded under non-IID data. This challenge is particularly severe in natural language processing, where client data may differ in linguistic style, topic distribution, and label semantics. Recent federated prompt-based methods address this issue by adapting large pre-trained language models (PLMs) through learnable prompts. However, most existing approaches rely on a single prompt generator, which limits their ability to capture diverse client characteristics. In this paper, we propose a novel mixture-of-prompt experts framework for federated text classification in non-IID settings. Our framework replaces the single dynamic prompt generator with multiple prompt experts that produce input-conditional prompts, and employs a gating network to dynamically combine expert outputs for each input instance. To balance generalization and personalization, the prompt experts are globally shared and aggregated via federated averaging, while the gating network is maintained locally on each client to enable client-specific expert selection without additional communication cost. This design allows the model to adapt at both the sample and client levels. We evaluate the proposed framework on the Stanford Sentiment Treebank, IMDB, and Banking77 benchmarks, spanning both binary sentiment classification and multi-class intent detection. Experimental results demonstrate that our model outperforms other baselines. Moreover, to comprehensively evaluate the proposed framework, we conduct sensitivity experiments across various hyperparameter configurations. These analyses demonstrate the stability of our model and provide insights into the trade-offs between performance and computational efficiency.

Scientific Reports
New York State Department of Health (US), Tunghai University (TW), Chi Mei Medical Center (TW), Taipei Medical University (TW)
Openalex Percentile: Top 9%
Privacy-Preserving Technologies in Data
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.