The unfinished business of evidence strength in software engineering: Current practices and future directions
摘要
Almost 20 years ago, the Software Engineering research community was introduced to the concept of evidence strength: the extent of confidence that the estimates of the effect of an intervention are correct. The motivation for assessing the strength of evidence is very practical: if one has high confidence in the beneficial effects of an intervention and its low harm, one can make strong recommendations for its adoption in the daily work of practitioners. Since then, software engineering researchers have started using GRADE (Grades of Recommendation, Assessment, Development, and Evaluation) as an approach to the assessment of strength of evidence. However, to the present day, no study has evaluated the extent to which systematic literature reviews in software engineering have effectively assessed the strength of evidence supporting their review findings.
Objective:We aim to understand the current strength assessment of evidence in software engineering secondary studies, answering questions about the most widely used methods, how researchers use or adapt them, and whether software engineering researchers make recommendations to practitioners from their review findings.
Method:We executed a tertiary study using automated and snowballing search procedures to accomplish our goal. For data extraction and analysis, we selected 22 papers in which the authors claimed to perform a strength assessment of evidence.
Results:GRADE is the most used method, although not necessarily appropriate considering the type of review. Additionally, software engineering researchers are not benefiting from its more recent versions or GRADE-CERQual (Confidence in the Evidence from Reviews of Qualitative research), a method specific to qualitative reviews. We also found that most strength assessments are generic for all review findings instead of specific for each review finding or outcome of interest, as the GRADE and GRADE-CERQual approaches recommend. Furthermore, very few reviews make explicit recommendations to practitioners based on their review findings, and none of them evaluated the strength of such recommendations.
Conclusions:A fundamental issue our study brings forward is that if the Software Engineering community decides to apply the original methods “as is” there is plenty of room for improvement in their faithful adoption. However, more than just adopting existing frameworks blindly—and at times incorrectly—it is important to reflect on what kinds of evidence our community produces and how best to assess it given the particularities of our field. Adapting existing methods could represent a step forward, yet SE researchers might also need to consider developing new ones from scratch.