INTERPRETABLE MACHINE LEARNING FOR PREDICTIVE HEPATOTOXICITY ASSESSMENT USING MOLECULAR FINGERPRINTS AND SHAP-BASED STRUCTURAL ANALYSIS

Authors

  • Ramesh Khanna Aluru Author
  • Narayan Padmaja Author
  • Busetty Pavan Kumar Author
  • Pullela Venkata Bala Annapurna Author
  • Poornima R Author
  • Balakonda Reddy Author
  • Maneesha LLS Author
  • Srujan Manohar MVN Author

DOI:

https://doi.org/10.4238/1n1q6n71

Keywords:

Drug-induced liver injury (DILI); Hepatotoxicity prediction; Quantitative structure–activity relationship (QSAR); XGBoost; SHAP; Molecular fingerprints; Computational toxicology; Structural alerts; Toxicity prediction.

Abstract

Drug-induced liver injury (DILI) remains one of the primary causes of drug attrition during preclinical and clinical development, emphasizing the need for accurate and interpretable computational toxicity prediction. This study presents an explainable quantitative structure–activity relationship (QSAR)-based machine learning framework for early hepatotoxicity assessment using molecular fingerprints and structural toxicity signatures. Publicly available HepG2 cytotoxicity data from PubChem BioAssay and oxidative stress response data from the Tox21 SR-ARE assay were employed to develop predictive models based on Morgan fingerprint representations of chemical structures. Three supervised machine learning algorithms, namely Random Forest, Extreme Gradient Boosting (XGBoost), and Deep Neural Networks, were systematically evaluated and compared using standard classification metrics. Among the investigated models, XGBoost demonstrated the highest predictive performance on the HepG2 dataset, achieving an area under the receiver operating characteristic curve (AUC-ROC) of 0.9555, indicating excellent discrimination between hepatotoxic and non-hepatotoxic compounds. Cross-assay validation using the Tox21 SR-ARE dataset resulted in comparatively lower predictive accuracy, highlighting the influence of assay-specific biological endpoints on model generalizability. To enhance model transparency, SHapley Additive exPlanations (SHAP) analysis was performed to identify molecular substructures contributing to toxicity predictions. The analysis revealed several recurring structural toxicity signatures, including nitrogen-containing heterocyclic systems, electron-withdrawing aromatic substituents, nitrogen-rich aromatic fragments, and sulfur-containing heterocycles such as thiazole and oxazole derivatives. Furthermore, a prototype HepatoRisk Index was developed by integrating prediction probabilities with structural toxicity signatures to facilitate compound prioritization during early-stage drug discovery. The proposed framework offers a transparent, scalable, and computationally efficient strategy for preliminary hepatotoxicity assessment, demonstrating the potential of explainable machine learning to support mechanism-informed drug safety evaluation and rational decision-making in computational toxicology.

Downloads

Published

2026-04-02

Issue

Section

Articles