Software categorization is a vital process involving the systematic grouping of software based on different criteria including the approach used and domain. This categorization has traditionally played a crucial role in software maintenance, aiding developers in locating specific programs, identifying their characteristics, and finding similar ones within extensive code repositories. However, manual categorization is known to be expensive, tedious, and time-consuming, highlighting the growing importance of automated categorization techniques. The objective is to develop a classifier capable of accurately and efficiently categorizing these repositories into predefined domains with less data. The contribution of this research is twofold. Firstly, we propose a model that enables the categorization of software repositories in terms of domains by using repository’s features, even with a limited amount of training data. Secondly, we conduct a comprehensive empirical evaluation to assess the impact of repository features and data augmentation technique on the task. This research aims to advance the field of software categorization, ultimately facilitating improved utilization of software repositories across diverse research domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Categorization of Software Repository Domains with Minimal Resources

  • Abdelhalim Hafedh Dahou,
  • Brigitte Mathiak

摘要

Software categorization is a vital process involving the systematic grouping of software based on different criteria including the approach used and domain. This categorization has traditionally played a crucial role in software maintenance, aiding developers in locating specific programs, identifying their characteristics, and finding similar ones within extensive code repositories. However, manual categorization is known to be expensive, tedious, and time-consuming, highlighting the growing importance of automated categorization techniques. The objective is to develop a classifier capable of accurately and efficiently categorizing these repositories into predefined domains with less data. The contribution of this research is twofold. Firstly, we propose a model that enables the categorization of software repositories in terms of domains by using repository’s features, even with a limited amount of training data. Secondly, we conduct a comprehensive empirical evaluation to assess the impact of repository features and data augmentation technique on the task. This research aims to advance the field of software categorization, ultimately facilitating improved utilization of software repositories across diverse research domains.