Stack Overflow, a widely used Internet platform for asking coding doubts, learning coding and sharing knowledge, facilitates public engagement through its question-answering forum. Users on Stack Overflow are required to manually enter a minimum of one tag for the questions they post. Tagging involves associating keywords with the sentence or question, aiding in categorization and accessibility. Analysis of tags on the platform revealed that a significant number of questions are labeled with multiple tags, and inaccuracies in tagging are common. This situation hinders users in searching for suitable tags. The primary objective of this research is to investigate, analyze and compare methods for developing an auto-tagging system using machine learning and deep learning techniques, along with Natural Language Processing-based data preprocessing steps. We have suggested robust and well-structured text preprocessing pipeline and addresses tag prediction as a multi-label classification problem unlike earlier methods where it was considered binary classification. The outcomes of the study yielded satisfactory results with opportunities for improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Machine Learning Algorithms in Automatic Question Tagging

  • Medha Joshi,
  • Ritika Kumari

摘要

Stack Overflow, a widely used Internet platform for asking coding doubts, learning coding and sharing knowledge, facilitates public engagement through its question-answering forum. Users on Stack Overflow are required to manually enter a minimum of one tag for the questions they post. Tagging involves associating keywords with the sentence or question, aiding in categorization and accessibility. Analysis of tags on the platform revealed that a significant number of questions are labeled with multiple tags, and inaccuracies in tagging are common. This situation hinders users in searching for suitable tags. The primary objective of this research is to investigate, analyze and compare methods for developing an auto-tagging system using machine learning and deep learning techniques, along with Natural Language Processing-based data preprocessing steps. We have suggested robust and well-structured text preprocessing pipeline and addresses tag prediction as a multi-label classification problem unlike earlier methods where it was considered binary classification. The outcomes of the study yielded satisfactory results with opportunities for improvement.