<p>Abstractive text summarization aims to generate semantically rich as well as coherent summaries that go beyond simple sentence extraction. This study introduces a novel architecture, dynamic multi-head attention-based LSTM (DMHA-LSTM), which enhances sequential modeling through an adaptive attention mechanism. Unlike Transformer-based architectures, the proposed model omits self-attention blocks and positional encodings, instead incorporating embedding dynamic attention directly within an LSTM encoder–decoder framework to improve contextual alignment and summary relevance. We evaluate the model on three benchmark datasets—Amazon Fine Food Reviews, BBC News, and CNN/DailyMail—across four train-test splits (60:40 to 90:10) to assess generalization. The proposed DMHA-LSTM consistently outperforms baseline models, including vanilla LSTM, static attention-LSTM, and Transformer-based models like PEGASUS, BART, and T5. It achieves peak scores of ROUGE-1: 0.46, ROUGE-2: 0.30, ROUGE-L: 0.45, BLEU: 0.25, and METEOR: 0.39. BLEU and METEOR scores range from 0.17 to 0.25 and 0.30–0.39 across datasets, demonstrating strong semantic consistency and lexical diversity. Furthermore, an ablation study is conducted an ablation investigation to quantify the contribution of the dynamic attention module, alongside a human evaluation measuring fluency and informativeness. These data confirm the model’s effectiveness in generating high-quality, linguistically coherent summaries. This work provides a robust, domain-independent framework for abstractive summarization and paves the way for future integration with pre-trained language models and deployment in low-resource environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimized dynamic multi-head attention for abstractive text summarization

  • Shafiya Mushtaq,
  • K. Veningston

摘要

Abstractive text summarization aims to generate semantically rich as well as coherent summaries that go beyond simple sentence extraction. This study introduces a novel architecture, dynamic multi-head attention-based LSTM (DMHA-LSTM), which enhances sequential modeling through an adaptive attention mechanism. Unlike Transformer-based architectures, the proposed model omits self-attention blocks and positional encodings, instead incorporating embedding dynamic attention directly within an LSTM encoder–decoder framework to improve contextual alignment and summary relevance. We evaluate the model on three benchmark datasets—Amazon Fine Food Reviews, BBC News, and CNN/DailyMail—across four train-test splits (60:40 to 90:10) to assess generalization. The proposed DMHA-LSTM consistently outperforms baseline models, including vanilla LSTM, static attention-LSTM, and Transformer-based models like PEGASUS, BART, and T5. It achieves peak scores of ROUGE-1: 0.46, ROUGE-2: 0.30, ROUGE-L: 0.45, BLEU: 0.25, and METEOR: 0.39. BLEU and METEOR scores range from 0.17 to 0.25 and 0.30–0.39 across datasets, demonstrating strong semantic consistency and lexical diversity. Furthermore, an ablation study is conducted an ablation investigation to quantify the contribution of the dynamic attention module, alongside a human evaluation measuring fluency and informativeness. These data confirm the model’s effectiveness in generating high-quality, linguistically coherent summaries. This work provides a robust, domain-independent framework for abstractive summarization and paves the way for future integration with pre-trained language models and deployment in low-resource environments.