Text classification in natural language processing (NLP) is advancing quickly, especially with the rise of transformer-based methods and large language models (LLMs). The daily generation of massive amounts of data presents significant challenges related to big data, particularly in text mining and classification. Deep learning-based text categorization is a considerable area of research with numerous applications, including sentiment analysis, company reports, news categorization, spam detection, modeling topics, stocks-related reports, question answering, and identifying hate speech. However, accurately extracting this information is still a big issue due to the large amount of textual data. This paper offers a comprehensive review of text classification approaches to emphasize the contrasts between typical approaches and newer, deep learning-based approaches, concentrating on their operational mechanisms and impact on input data transformation. The study employs NLP-driven methods like co-citation and bibliographic coupling in conjunction with conventional research methodologies. In order to facilitate the researcher, this paper provides a systematic review of various deep learning approaches such as CNN, LSTM, GRU, BiLSTM, and BiGRU and LLMs including, BERT, GPT-3.5, GPT-4, RoBERTa, Flan-T5, and XLNet for text classification, highlighting their strengths and limitations. In conclusion, we emphasizing the new insights of the main implications, outlined potential directions for future research, and highlighted the challenges currently being faced in this particular research field. It also aims to familiarize readers with the different subtasks and relevant literature in the text classification process. We hope our discussion will encourage readers to seek out new and improved techniques for text classification that can be applied across various domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Survey on Text Classification Using Deep Learning Approaches

  • Dhurgham Ali Mohammed Alhasani,
  • Kalyani A. Patel

摘要

Text classification in natural language processing (NLP) is advancing quickly, especially with the rise of transformer-based methods and large language models (LLMs). The daily generation of massive amounts of data presents significant challenges related to big data, particularly in text mining and classification. Deep learning-based text categorization is a considerable area of research with numerous applications, including sentiment analysis, company reports, news categorization, spam detection, modeling topics, stocks-related reports, question answering, and identifying hate speech. However, accurately extracting this information is still a big issue due to the large amount of textual data. This paper offers a comprehensive review of text classification approaches to emphasize the contrasts between typical approaches and newer, deep learning-based approaches, concentrating on their operational mechanisms and impact on input data transformation. The study employs NLP-driven methods like co-citation and bibliographic coupling in conjunction with conventional research methodologies. In order to facilitate the researcher, this paper provides a systematic review of various deep learning approaches such as CNN, LSTM, GRU, BiLSTM, and BiGRU and LLMs including, BERT, GPT-3.5, GPT-4, RoBERTa, Flan-T5, and XLNet for text classification, highlighting their strengths and limitations. In conclusion, we emphasizing the new insights of the main implications, outlined potential directions for future research, and highlighted the challenges currently being faced in this particular research field. It also aims to familiarize readers with the different subtasks and relevant literature in the text classification process. We hope our discussion will encourage readers to seek out new and improved techniques for text classification that can be applied across various domains.