<p>Accurate prediction of <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\mathrm {CO_2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">CO</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> solubility in physical solvents is crucial for advancing carbon capture technologies. However, although existing machine learning and deep learning models offer high accuracy, their “black-box" nature limits their application in engineering practice and scientific discovery. To address the contradiction between accuracy and interpretability, this paper applies Kolmogorov–Arnold Networks (KAN) to <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\mathrm {CO_2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">CO</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> solubility prediction for the first time, establishing a systematic framework that spans from performance validation to model simplification. First, in terms of the full feature set, KAN performs noticeably better than 10 mainstream methods. Subsequently, to mitigate the impact of feature redundancy on model generalization, seven feature selection methods are systematically evaluated, and Lasso regression is ultimately adopted to determine a 4-dimensional optimal feature subset. The KAN constructed based on this subset achieves a coefficient of determination <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\({(R^2)}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo stretchy="false">(</mo> <msup> <mi>R</mi> <mn>2</mn> </msup> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> of 0.999115. Compared to the second-best method (XGB), <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\({R^2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>R</mi> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> is improved by 1.17%. More importantly, through pruning and symbolization, KAN yields mathematical expressions highly consistent with physicochemical mechanisms and maintains excellent accuracy (<InlineEquation ID="IEq7"> <EquationSource Format="TEX">\({R^2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>R</mi> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> of 0.941937), and achieves an 8.18-fold inference speedup over the original KAN model. This work provides interpretable theoretical guidance for solvent selection and optimization in <InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(\mathrm {CO_2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="normal">CO</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> capture processes, contributing to the advancement of carbon capture technologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interpretable Prediction of \(\mathrm {CO_2}\) Solubility in Physical Solvents Using Kolmogorov–Arnold Networks

  • Yu Hua,
  • Kuanping Gong,
  • Yongquan Jiang,
  • Fuhua Li

摘要

Accurate prediction of \(\mathrm {CO_2}\) CO 2 solubility in physical solvents is crucial for advancing carbon capture technologies. However, although existing machine learning and deep learning models offer high accuracy, their “black-box" nature limits their application in engineering practice and scientific discovery. To address the contradiction between accuracy and interpretability, this paper applies Kolmogorov–Arnold Networks (KAN) to \(\mathrm {CO_2}\) CO 2 solubility prediction for the first time, establishing a systematic framework that spans from performance validation to model simplification. First, in terms of the full feature set, KAN performs noticeably better than 10 mainstream methods. Subsequently, to mitigate the impact of feature redundancy on model generalization, seven feature selection methods are systematically evaluated, and Lasso regression is ultimately adopted to determine a 4-dimensional optimal feature subset. The KAN constructed based on this subset achieves a coefficient of determination \({(R^2)}\) ( R 2 ) of 0.999115. Compared to the second-best method (XGB), \({R^2}\) R 2 is improved by 1.17%. More importantly, through pruning and symbolization, KAN yields mathematical expressions highly consistent with physicochemical mechanisms and maintains excellent accuracy ( \({R^2}\) R 2 of 0.941937), and achieves an 8.18-fold inference speedup over the original KAN model. This work provides interpretable theoretical guidance for solvent selection and optimization in \(\mathrm {CO_2}\) CO 2 capture processes, contributing to the advancement of carbon capture technologies.