Efficient identification of persistent, bioaccumulative, and toxic (PBT) chemicals is crucial for chemical risk management. However, established machine learning or QSAR models often rely on descriptors or molecular graph, facing inherent trade-offs among predictive performance, interpretability, and computational efficiency. In this study, a computationally efficient and interpretable framework was developed based on a curated dataset containing 1,749 confirmed PBT chemicals and 2,775 non-PBT chemicals, compiled from regulatory inventories, peer-reviewed literature, and experimental records. Twenty-four descriptors and one hundred and sixty-six MACCS fingerprints were used to represent molecular structures. Nine widely used algorithms were systematically evaluated under a fixed data split and hyperparameter-tuning protocol, reporting multi-metric predictive performance together with quantified compute overhead (training/tuning and inference burden). CatBoost achieved the best performance with generalization F1-score of 0.9153 and was selected for subsequent reliability and interpretability analyses. SHAP analysis identified halogenation, molecular complexity, and lipophilicity as dominant decision drivers in the model for PBT classification and SHAP interaction analysis indicated non-additive dependencies among these drivers. A similarity-defined applicability domain (AD) and activity-cliff flagging were used to mark low-confidence subdomains for expert review or additional evidence assessment in external screening. The framework was also benchmarked against the OECD QSAR Toolbox, evaluating strict coverage and performance on matched comparable subsets. Overall, the proposed framework balanced predictive accuracy, interpretability, and computational efficiency, providing a practical tool for high-throughput PBT screening.
Yan et al. (Wed,) studied this question.