Evaluating the Effectiveness of Large Language Models in Multi-Document Summarization of Bangla News Articles
摘要
Multi-Document Summarization (MDS) in low-resource languages like Bangla faces challenges due to limited datasets, tools, and benchmarks. This study evaluates eight open-source Large Language Models (LLMs), including LLaMA, Gemma, DeepSeek, and Mistral variants, using the BUSUM-BNLP dataset and a prompt-based summarization framework. Model outputs are assessed with ROUGE, BLEU, and BERTScore metrics. Mistral-8x22B achieves the highest ROUGE-1 (28.68), BLEU-4 (5.71), and BERTScore-F1 (0.7454), indicating strong lexical and semantic performance. Despite being lightweight, DeepSeek-v3-Base scores highest on ROUGE-2 and ROUGE-L, showing better structural coherence. LLaMA-3.2-1B performs poorly, averaging only 29 words per summary. BanglaBERT, though not generative, outperforms Gemma2-9B-IT and LLaMA-3.2-1B in some cases, underscoring the value of language-specific pretraining. A small-scale human evaluation of fluency and informativeness aligns with metric-based results, confirming Mistral and DeepSeek’s superior performance. These findings highlight the potential and limitations of open-source LLMs for Bangla MDS and offer benchmarks for future research in low-resource summarization.