Identifying high-quality research topics in R&D organizations using explainable machine learning
摘要
Scientific research and development are significant drivers of innovative organizations. The selection of high-quality research topics to invest in directly contributes to the prospects of R&D-based organizations. The emergence and evolution of the research topic is a complex process involving nonlinear interactions among several factors. Machine learning approaches are adept at fitting high-dimensional complex interactions, and we selected the eight classifiers to identify high-quality research topics. By training these models using high-quality research topic datasets, we found that the Gradient Boosting classifier exhibits superior performance in detecting high-quality research topics, as evidenced by its F1-score of 94.45%. Further, we conducted SHapley Additive exPlanations (SHAP) analysis to interpret the model’s decision-making process. Our findings demonstrate that high-quality research topics exhibit three distinctive signatures: (1) rapid engagement by the scientific community, evidenced by a sharp rise in publications and citation rates within 2–3 years of emergence; (2) disproportionate early adoption by highly cited authors; and (3) a characteristic network topology featuring a sparse, ego-centered structure with limited direct connections to existing topics and weak interlinkages among neighboring domains.