Document classification is one of the mostly studied topics in natural language processing (NLP), which aims to classify documents into given classes. Although previous studies focused on using supervised or deep learning models, such as SVM and CNN, recent advances in large language models (LLMs) introduce new opportunities to improve document classification accuracy. However, few works have addressed the different performance of traditional models and LLMs when applied to document classification. In this paper, we performed a comparative experimental study on various datasets to evaluate different LLMs and traditional models for document classification. The unique contributions of the paper are two-fold. First, we pre-sent a systematic experimental study to compare ten models in document classification. These models cover four traditional supervised models, two deep learning models, and four LLMs with various parameter sizes, providing a relatively complete framework to reveal the performance of different models in document classification. Second, we investigate the fine-tuning approach and Mixture of Experts (MoE) for LLMs and experimentally study the performance of fine-tuned and MoE-based LLM (DeepSeek V3). The results show that the original LLMs, without fine-tuning, do not perform better than traditional supervised and deep learning models in document classification if the knowledge base of LLMs does not cover the test dataset in their training. In addition, LLMs’ fine-tuning and MoE architecture can help improve accuracy, outperforming the original LLMs without fine-tuning. Finally, we present some insights and suggestions for future research on LLM-based document classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Will Large Language Models Outperform Traditional Models in Document Classification?

  • Haoying Jin

摘要

Document classification is one of the mostly studied topics in natural language processing (NLP), which aims to classify documents into given classes. Although previous studies focused on using supervised or deep learning models, such as SVM and CNN, recent advances in large language models (LLMs) introduce new opportunities to improve document classification accuracy. However, few works have addressed the different performance of traditional models and LLMs when applied to document classification. In this paper, we performed a comparative experimental study on various datasets to evaluate different LLMs and traditional models for document classification. The unique contributions of the paper are two-fold. First, we pre-sent a systematic experimental study to compare ten models in document classification. These models cover four traditional supervised models, two deep learning models, and four LLMs with various parameter sizes, providing a relatively complete framework to reveal the performance of different models in document classification. Second, we investigate the fine-tuning approach and Mixture of Experts (MoE) for LLMs and experimentally study the performance of fine-tuned and MoE-based LLM (DeepSeek V3). The results show that the original LLMs, without fine-tuning, do not perform better than traditional supervised and deep learning models in document classification if the knowledge base of LLMs does not cover the test dataset in their training. In addition, LLMs’ fine-tuning and MoE architecture can help improve accuracy, outperforming the original LLMs without fine-tuning. Finally, we present some insights and suggestions for future research on LLM-based document classification.