<p>This study provides a decade-long and pioneering systematic review of data envelopment analysis (DEA) integrated with machine learning (ML), consolidating methodological advances, empirical patterns, and emerging learning-based frontier models. Following the PRISMA protocol, 521 records were initially identified, 313 eligible studies were retrieved, and 210 peer-reviewed articles were systematically analyzed, complemented by 30 frontier studies on learning-based models. The findings show rapid growth in DEA-ML research, with annual publications rising from fewer than 20 studies during 2015–2019 to 70 studies in 2025. Classical DEA models remain dominant, with CCR and BCC appearing in 73 and 71 studies, respectively. On the ML side, Artificial Neural Networks (ANNs) dominate regression-based applications, with 91 studies, followed by Support Vector Regression (SVR) with 28 studies and Random Forests (RFs) with 27 studies. Software reporting remains incomplete, as only 152 of the 210 studies explicitly reported computational environments; among these, <i>MATLAB</i> was most frequent, with 50 studies, while <i>R</i> and <i>Python</i> each appeared in 35 studies. Domain analysis shows concentration in banking &amp; finance, healthcare, and environment &amp; sustainability, while insurance, construction, and forestry remain underexplored. Beyond traditional hybrids, this review synthesizes learning-based frontier models, such as Efficiency Analysis Trees (EAT) and its variants, clarifying how emerging statistical learning can support frontier estimation. The study identifies persistent gaps in small-data learning, interpretability, benchmarking, and software integration, and later proposes a <i>Python</i>-based implementation pathway and a reproducibility-oriented research agenda. The findings offer actionable insights for researchers, practitioners, software developers, and policymakers who require interpretable efficiency analytics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Static Efficiency to Predictive Intelligence: A Review of Data Envelopment Analysis Integrated with Machine Learning

  • Temitope Olubanjo Kehinde,
  • Waqar Ahmed Khan,
  • Sai-Ho Chung,
  • Kelvin K. Orisaremi,
  • Joseph Akpan,
  • Sujan Piya

摘要

This study provides a decade-long and pioneering systematic review of data envelopment analysis (DEA) integrated with machine learning (ML), consolidating methodological advances, empirical patterns, and emerging learning-based frontier models. Following the PRISMA protocol, 521 records were initially identified, 313 eligible studies were retrieved, and 210 peer-reviewed articles were systematically analyzed, complemented by 30 frontier studies on learning-based models. The findings show rapid growth in DEA-ML research, with annual publications rising from fewer than 20 studies during 2015–2019 to 70 studies in 2025. Classical DEA models remain dominant, with CCR and BCC appearing in 73 and 71 studies, respectively. On the ML side, Artificial Neural Networks (ANNs) dominate regression-based applications, with 91 studies, followed by Support Vector Regression (SVR) with 28 studies and Random Forests (RFs) with 27 studies. Software reporting remains incomplete, as only 152 of the 210 studies explicitly reported computational environments; among these, MATLAB was most frequent, with 50 studies, while R and Python each appeared in 35 studies. Domain analysis shows concentration in banking & finance, healthcare, and environment & sustainability, while insurance, construction, and forestry remain underexplored. Beyond traditional hybrids, this review synthesizes learning-based frontier models, such as Efficiency Analysis Trees (EAT) and its variants, clarifying how emerging statistical learning can support frontier estimation. The study identifies persistent gaps in small-data learning, interpretability, benchmarking, and software integration, and later proposes a Python-based implementation pathway and a reproducibility-oriented research agenda. The findings offer actionable insights for researchers, practitioners, software developers, and policymakers who require interpretable efficiency analytics.