<p>The accurate prediction of exhaust emissions from compression-ignition engines fueled with alternative biofuels remains a long-standing challenge for sustainable engine development, particularly when experimental datasets are small. This study presents a systematic comparison of six machine learning (ML) models linear regression (LR), polynomial regression (degree 2), support vector regression with radial basis function kernel (SVR-RBF), random forest (RF), gradient boosting (GB), and an artificial neural network (ANN-MLP) for predicting six emission parameters (CO, HC, CO<sub>2</sub>, O<sub>2</sub>, NO<sub>x</sub>, and the air–fuel equivalence ratio λ) of a single-cylinder variable compression ratio diesel engine. The engine was operated with three neat fuels (conventional diesel, rubber seed oil biodiesel produced by two-stage acid–base transesterification, and Chlorella vulgaris microalgae biodiesel) at three compression ratios (16:1, 17:1, 18:1) and five loads (0–100% in 25% increments), forming a balanced 45-condition factorial. All models were assessed using leave-one-out cross-validation (LOOCV) with fixed literature-based hyperparameters. Gradient boosting achieved the highest mean R<sup>2</sup> (0.816) across the six outputs, followed by RF (0.787), Poly (0.756), LR (0.713), SVR-RBF (0.496), and ANN-MLP (− 1.758). GB topped four of six outputs (λ, O<sub>2</sub>, HC, CO); RF led on NO<sub>x</sub> (R<sup>2</sup> = 0.934), and Poly narrowly led on CO<sub>2</sub> (R<sup>2</sup> = 0.963 vs GB 0.960). The ANN-MLP yielded negative R<sup>2</sup> on four of six outputs, traced through architecture sensitivity sweeps and a direct LOOCV-versus-5-fold cross-validation comparison to an unfavorable parameter-to-sample ratio (~ 1.4:1) rather than to validation choice. The six emissions stratify into three predictability levels: highly predictable (λ, O<sub>2</sub>, CO<sub>2</sub>; R<sup>2</sup> &gt; 0.96), moderately predictable (NO<sub>x</sub>, HC; R<sup>2</sup> = 0.70–0.93), and poorly predictable (CO; R<sup>2</sup> &lt; 0.40). For the present 45-sample regime, gradient boosting is recommended as the model of choice, and neural networks should be applied cautiously to small multi-fuel emission datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative evaluation of six machine learning models for multi-fuel variable compression ratio diesel engine emission prediction under leave-one-out cross-validation

  • Vasanthaseelan Sathiyaseelan,
  • Arivazhagan Sampathkumar,
  • Senthilnathan Natarajan,
  • Bharathwaaj Ramani,
  • E. Joel,
  • V. Sakthi Murugan,
  • P. Manoj Kumar,
  • Dawit Tafesse Gebreyohannes

摘要

The accurate prediction of exhaust emissions from compression-ignition engines fueled with alternative biofuels remains a long-standing challenge for sustainable engine development, particularly when experimental datasets are small. This study presents a systematic comparison of six machine learning (ML) models linear regression (LR), polynomial regression (degree 2), support vector regression with radial basis function kernel (SVR-RBF), random forest (RF), gradient boosting (GB), and an artificial neural network (ANN-MLP) for predicting six emission parameters (CO, HC, CO2, O2, NOx, and the air–fuel equivalence ratio λ) of a single-cylinder variable compression ratio diesel engine. The engine was operated with three neat fuels (conventional diesel, rubber seed oil biodiesel produced by two-stage acid–base transesterification, and Chlorella vulgaris microalgae biodiesel) at three compression ratios (16:1, 17:1, 18:1) and five loads (0–100% in 25% increments), forming a balanced 45-condition factorial. All models were assessed using leave-one-out cross-validation (LOOCV) with fixed literature-based hyperparameters. Gradient boosting achieved the highest mean R2 (0.816) across the six outputs, followed by RF (0.787), Poly (0.756), LR (0.713), SVR-RBF (0.496), and ANN-MLP (− 1.758). GB topped four of six outputs (λ, O2, HC, CO); RF led on NOx (R2 = 0.934), and Poly narrowly led on CO2 (R2 = 0.963 vs GB 0.960). The ANN-MLP yielded negative R2 on four of six outputs, traced through architecture sensitivity sweeps and a direct LOOCV-versus-5-fold cross-validation comparison to an unfavorable parameter-to-sample ratio (~ 1.4:1) rather than to validation choice. The six emissions stratify into three predictability levels: highly predictable (λ, O2, CO2; R2 > 0.96), moderately predictable (NOx, HC; R2 = 0.70–0.93), and poorly predictable (CO; R2 < 0.40). For the present 45-sample regime, gradient boosting is recommended as the model of choice, and neural networks should be applied cautiously to small multi-fuel emission datasets.