
Anusree Roy, Fahrin Hossain Sunaira, Tateyama Orpa, Farasha Shamma Yussouf, Intisar Tahmid Naheen, Riasat Khan · PLoS ONE 2026 · 2026
DOI: 10.1371/journal.pone.0357890
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This work aims to develop a domain-adapted Bengali text summarization model by training and fine-tuning on general and domain-specific datasets with categories such as state, international, and sports. Flan-T5 and mT5 LLMs were trained on the XLSUM Bengali dataset for text summarization, and they were trained on source domains and fine-tuned on the target domains for domain adaptation. In general-domain summarization on the XLSum Bengali dataset, mT5 achieved stronger overall performance than Flan-T5, with ROUGE-1 and BERTScore values of 0.21 and 0.72, respectively. The Flan-T5-XLSUM model outperformed other LLMs, achieving a ROUGE score of 0.79, a BLEU score of 0.13 and a BERTScore of 0.86. A human evaluation involving 28 participants was conducted to assess summary fluency, adequacy, and factual consistency across domains, confirming the qualitative effectiveness of the proposed models. A custom dataset is developed with manually annotated categories for domain adaptation purposes, contributing to domain-specific Bengali data. The ablation study illustrated that pre-training and fine-tuning on domain-specific data significantly enhanced model performance, with Flan-T5 fine-tuned on sports data. Three explainability methods (LIME, SHAP, and BertViz) revealed that geographically and contextually meaningful tokens strongly influenced the domain-adaptive Flan-T5 Bengali summarization model. The adapter-based Flan-T5 architecture enables effective multi-domain fine-tuning with less than 1% of trainable parameters while preserving model performance.
No comments yet — start the discussion below.