<p>Solvents play a critical role in separation processes by selectively dissolving or extracting specific components from a mixture, enabling their effective separation. The choice of solvent influences the efficiency, selectivity, and energy consumption of the process, making it a key factor in optimizing separation techniques such as distillation, extraction, and crystallization. The present study highlights a development of several machine learning (ML) models to predict the solubility of a solute in a solvent by using their SMILES as inputs. Molecular descriptors of solutes and solvents are obtained from SMILES of the original dataset. The top 5 descriptors are selected based on Pearson’s coefficient for solutes and solvents and are considered as inputs along with temperature. Different ML models are used for solubility prediction including linear models (linear, lasso, and ridge regression models), tree-based models (decision tree, random forest regressor, gradient boost, xgboost models and AdaBoost models) and other models (support vector regressor, k-nearest neighbor). The random forest model performed well with R<sup>2</sup> = 0.98, RMSE = 0.0121, and MSE = 0.0001 using training dataset, R<sup>2</sup> = 0.95, RMSE = 0.0266, and MSE = 0.0007 using testing dataset, and R<sup>2</sup> = 0.97, RMSE = 0.0161, and MSE = 0.0003 with the overall data. The prediction capability of the model is analyzed with respect to different descriptors and with respect to solutes and solvents, and with respect to temperature dependency. The model selected in the present study can be directly used for solvent design in various separation processes.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Models for Estimation of Solubility for A Wide Range of Solutes in Multiple Solvents Using Molecular Descriptors

  • Shriya Deshpande,
  • K. Yamuna Rani

摘要

Solvents play a critical role in separation processes by selectively dissolving or extracting specific components from a mixture, enabling their effective separation. The choice of solvent influences the efficiency, selectivity, and energy consumption of the process, making it a key factor in optimizing separation techniques such as distillation, extraction, and crystallization. The present study highlights a development of several machine learning (ML) models to predict the solubility of a solute in a solvent by using their SMILES as inputs. Molecular descriptors of solutes and solvents are obtained from SMILES of the original dataset. The top 5 descriptors are selected based on Pearson’s coefficient for solutes and solvents and are considered as inputs along with temperature. Different ML models are used for solubility prediction including linear models (linear, lasso, and ridge regression models), tree-based models (decision tree, random forest regressor, gradient boost, xgboost models and AdaBoost models) and other models (support vector regressor, k-nearest neighbor). The random forest model performed well with R2 = 0.98, RMSE = 0.0121, and MSE = 0.0001 using training dataset, R2 = 0.95, RMSE = 0.0266, and MSE = 0.0007 using testing dataset, and R2 = 0.97, RMSE = 0.0161, and MSE = 0.0003 with the overall data. The prediction capability of the model is analyzed with respect to different descriptors and with respect to solutes and solvents, and with respect to temperature dependency. The model selected in the present study can be directly used for solvent design in various separation processes.