Document categorization in Bengali involves automatically classifying documents based on keywords, with effective keyword selection being crucial. Popular methods include TF-IDF, Naive Bayes, and KNN. This research introduces the novel use of Term Graph Models (TGM) in Bengali document categorization, exploring its integration with existing models. The study experiments with various feature selection methods, including subsets of up to three features, and aims to improve accuracy and reduce space complexity by occasionally removing less important features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploration of Machine Learning Technique for Bengali Textual Content Classification

  • Arnab Maitra,
  • Soubhik Ghosh,
  • Dhiman Roy,
  • Vedatrayee Chatterjee

摘要

Document categorization in Bengali involves automatically classifying documents based on keywords, with effective keyword selection being crucial. Popular methods include TF-IDF, Naive Bayes, and KNN. This research introduces the novel use of Term Graph Models (TGM) in Bengali document categorization, exploring its integration with existing models. The study experiments with various feature selection methods, including subsets of up to three features, and aims to improve accuracy and reduce space complexity by occasionally removing less important features.