Domain-adaptive Bengali text summarization: Harnessing transformer-based large language models in closed-set contexts

This work aims to develop a domain-adapted Bengali text summarization model by training and fine-tuning on general and domain-specific datasets with categories such as state, international, and sports. Flan-T5 and mT5 LLMs were trained on the XLSUM Bengali dataset for text summarization, and they were trained on source domains and fine-tuned on the target domains for domain adaptation. In general-domain summarization on the XLSum Bengali dataset, mT5 achieved stronger overall performance than Flan-T5, with ROUGE-1 and BERTScore values of 0.21 and 0.72, respectively. The Flan-T5-XLSUM model outperformed other LLMs, achieving a ROUGE score of 0.79, a BLEU score of 0.13 and a BERTScore of 0.86. A human evaluation involving 28 participants was conducted to assess summary fluency, adequacy, and factual consistency across domains, confirming the qualitative effectiveness of the proposed models. A custom dataset is developed with manually annotated categories for domain adaptation purposes, contributing to domain-specific Bengali data. The ablation study illustrated that pre-training and fine-tuning on domain-specific data significantly enhanced model performance, with Flan-T5 fine-tuned on sports data. Three explainability methods (LIME, SHAP, and BertViz) revealed that geographically and contextually meaningful tokens strongly influenced the domain-adaptive Flan-T5 Bengali summarization model. The adapter-based Flan-T5 architecture enables effective multi-domain fine-tuning with less than 1% of trainable parameters while preserving model performance.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-10-05
DOI
https://doi.org/10.1371/journal.pone.0357890
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Domain-adaptive Bengali text summarization: Harnessing transformer-based large language models in closed-set contexts

Riasat Khan, Intisar Tahmid Naheen, Anusree Roy, Fahrin Hossain Sunaira et al.
PLoS ONE
Topic Modeling
article

Domain-adaptive Bengali text summarization: Harnessing transformer-based large language models in closed-set contexts

Riasat Khan, Intisar Tahmid Naheen, Anusree Roy, Fahrin Hossain Sunaira, Tateyama Orpa, Farasha Shamma Yussouf
article en

Abstract

This work aims to develop a domain-adapted Bengali text summarization model by training and fine-tuning on general and domain-specific datasets with categories such as state, international, and sports. Flan-T5 and mT5 LLMs were trained on the XLSUM Bengali dataset for text summarization, and they were trained on source domains and fine-tuned on the target domains for domain adaptation. In general-domain summarization on the XLSum Bengali dataset, mT5 achieved stronger overall performance than Flan-T5, with ROUGE-1 and BERTScore values of 0.21 and 0.72, respectively. The Flan-T5-XLSUM model outperformed other LLMs, achieving a ROUGE score of 0.79, a BLEU score of 0.13 and a BERTScore of 0.86. A human evaluation involving 28 participants was conducted to assess summary fluency, adequacy, and factual consistency across domains, confirming the qualitative effectiveness of the proposed models. A custom dataset is developed with manually annotated categories for domain adaptation purposes, contributing to domain-specific Bengali data. The ablation study illustrated that pre-training and fine-tuning on domain-specific data significantly enhanced model performance, with Flan-T5 fine-tuned on sports data. Three explainability methods (LIME, SHAP, and BertViz) revealed that geographically and contextually meaningful tokens strongly influenced the domain-adaptive Flan-T5 Bengali summarization model. The adapter-based Flan-T5 architecture enables effective multi-domain fine-tuning with less than 1% of trainable parameters while preserving model performance.

PLoS ONEVol. 21(10)
North South University (BD)
Openalex Percentile: Top 10%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.