Predicting toxicity is essential for monitoring environmental health and evaluating pharmacological safety. To create machine learning models for toxicity classification, this study makes use of the extensive Toxic Exposome Database (T3DB), which includes 3,678 toxins with comprehensive chemical, toxicological, and biomolecular interaction data. Using extracted LD50 values and chemical descriptors, we assess several algorithms for toxicity level prediction, such as Random Forest, Gradient Boosting, and Artificial Neural Networks (ANN). Our findings show that when compared to deep learning approaches, ensemble methods—in particular, LightGBM—perform better, achieving 94.51% accuracy. The study highlights the relative advantages of various machine learning techniques in managing complex toxin data and offers a framework for automated toxicity classification using this abundant biochemical resource. These results have important ramifications for monitoring of public health, environmental risk assessment, and drug safety evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Machine Learning Study for LD50-Based Toxicity Classification and Prediction Using the T3DB Database

  • Sanjana Reddy Singam,
  • Aditya Anant,
  • Rudransh Vats,
  • Sahith Reddy Aleti,
  • C. O. Prakash

摘要

Predicting toxicity is essential for monitoring environmental health and evaluating pharmacological safety. To create machine learning models for toxicity classification, this study makes use of the extensive Toxic Exposome Database (T3DB), which includes 3,678 toxins with comprehensive chemical, toxicological, and biomolecular interaction data. Using extracted LD50 values and chemical descriptors, we assess several algorithms for toxicity level prediction, such as Random Forest, Gradient Boosting, and Artificial Neural Networks (ANN). Our findings show that when compared to deep learning approaches, ensemble methods—in particular, LightGBM—perform better, achieving 94.51% accuracy. The study highlights the relative advantages of various machine learning techniques in managing complex toxin data and offers a framework for automated toxicity classification using this abundant biochemical resource. These results have important ramifications for monitoring of public health, environmental risk assessment, and drug safety evaluation.