<p>Software requirements classification has been an imperative archetype in software engineering because it sets the foundation for the software engineering process. At first, classification of requirements was done manually, but more recently, machine learning, natural language processing, and generative artificial intelligence (AI) have facilitated the task. Software requirements with reduced ambiguity, misunderstanding, and development costs are always the prime goals, which can be achieved with these automated methods. However, optimization of the process of software requirements classification with improved preprocessing and data representations for the transformer-based models is still a demanding research area. The goals of the study are to optimize the software requirement classification process with upgraded accuracy and attainment of better precision, recall, and F1-scores for each class with the implementation of AI. The study proposes a fused back translation and BERT model, comprising a fusion of two models. In the first model, the imbalanced dataset is balanced using back translation through Marian. Samples of underrepresented classes in English are translated into German and then back to English to synthesize the new samples and increase the dataset. This augmented dataset is then preprocessed and fed into Bidirectional Encoder Representations from Transformer (BERT), which is used for the classification of software requirements. The proposed model has achieved an accuracy of 94.3%, precision of 0.91, recall of 0.89, and F1-score of 0.90. These significant results illustrate the efficacy of the proposed model in optimizing the software requirements classification, demonstrating the applications of AI in the field of software engineering.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Fused Back Translation and BERT Model for Optimizing Software Requirements Classification

  • Mehwish Naseer,
  • Usman Qamar,
  • Saad Almutairi

摘要

Software requirements classification has been an imperative archetype in software engineering because it sets the foundation for the software engineering process. At first, classification of requirements was done manually, but more recently, machine learning, natural language processing, and generative artificial intelligence (AI) have facilitated the task. Software requirements with reduced ambiguity, misunderstanding, and development costs are always the prime goals, which can be achieved with these automated methods. However, optimization of the process of software requirements classification with improved preprocessing and data representations for the transformer-based models is still a demanding research area. The goals of the study are to optimize the software requirement classification process with upgraded accuracy and attainment of better precision, recall, and F1-scores for each class with the implementation of AI. The study proposes a fused back translation and BERT model, comprising a fusion of two models. In the first model, the imbalanced dataset is balanced using back translation through Marian. Samples of underrepresented classes in English are translated into German and then back to English to synthesize the new samples and increase the dataset. This augmented dataset is then preprocessed and fed into Bidirectional Encoder Representations from Transformer (BERT), which is used for the classification of software requirements. The proposed model has achieved an accuracy of 94.3%, precision of 0.91, recall of 0.89, and F1-score of 0.90. These significant results illustrate the efficacy of the proposed model in optimizing the software requirements classification, demonstrating the applications of AI in the field of software engineering.