<p>Pharmaceutical pollutants are increasingly recognized as emerging contaminants in aquatic environments. Their persistence, bioactivity, and resistance to conventional treatment processes raise ecological and human health concerns, including the spread of antimicrobial resistance. Adsorption has emerged as a promising polishing step for their removal, but adsorption capacity (Qe, mg/g) varies widely depending on molecular structure and operational conditions, making predictive modeling essential. In this work, we developed machine learning models to predict adsorption capacities for Aspirin, Caffeine, Carbamazepine, Ketoprofen, Sulfamethoxazole, Nimesulide, and Paracetamol using chemoinformatics descriptors derived from SMILES strings and experimental inputs, including equilibrium concentration (Ce), initial concentration (C<sub>0</sub>), temperature, and contact time. Feature reduction with LassoCV and multicollinearity analysis yielded a compact, chemically interpretable descriptor set. Support Vector Regression (SVR), Extreme Gradient Boosting (XGB), and Artificial Neural Networks (ANN) were optimized with Optuna and evaluated using cross-validation. XGB delivered the best predictive performance (R2 = 0.997, RMSE = 2.62&#xa0;mg/g), outperforming SVR and ANN. SHAP analysis highlighted the influence of charge-partitioned surface areas and nitro functionalities on adsorption outcomes. The best-performing model was deployed in a Streamlit application, enabling predictions of Qe from SMILES and experimental conditions with built-in applicability-domain checks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting adsorption capacities of pharmaceutical pollutants using chemoinformatics and machine learning techniques

  • Hakim Bouzemlal,
  • Mohamed Hentabli,
  • Maamar Laidi,
  • Ykhlef Laidani,
  • Mohamed Kouider Amar,
  • Abdellah Ibrir,
  • Jie Zhang

摘要

Pharmaceutical pollutants are increasingly recognized as emerging contaminants in aquatic environments. Their persistence, bioactivity, and resistance to conventional treatment processes raise ecological and human health concerns, including the spread of antimicrobial resistance. Adsorption has emerged as a promising polishing step for their removal, but adsorption capacity (Qe, mg/g) varies widely depending on molecular structure and operational conditions, making predictive modeling essential. In this work, we developed machine learning models to predict adsorption capacities for Aspirin, Caffeine, Carbamazepine, Ketoprofen, Sulfamethoxazole, Nimesulide, and Paracetamol using chemoinformatics descriptors derived from SMILES strings and experimental inputs, including equilibrium concentration (Ce), initial concentration (C0), temperature, and contact time. Feature reduction with LassoCV and multicollinearity analysis yielded a compact, chemically interpretable descriptor set. Support Vector Regression (SVR), Extreme Gradient Boosting (XGB), and Artificial Neural Networks (ANN) were optimized with Optuna and evaluated using cross-validation. XGB delivered the best predictive performance (R2 = 0.997, RMSE = 2.62 mg/g), outperforming SVR and ANN. SHAP analysis highlighted the influence of charge-partitioned surface areas and nitro functionalities on adsorption outcomes. The best-performing model was deployed in a Streamlit application, enabling predictions of Qe from SMILES and experimental conditions with built-in applicability-domain checks.