Breast cancer remains a leading cause of cancer morbidity and mortality, motivating rapid discovery of target-focused inhibitors and chemical probes that modulate hormone biosynthesis and metastatic progression. Among clinically relevant molecular targets are aromatase (CYP19A1) and matrix metalloproteinases MMP-2 and MMP-9. In this work, we present a unified cheminformatics machine-learning framework for predicting potent small-molecule inhibitors of CYP19A1, MMP-2, and MMP-9 using curated ChEMBL bioactivity data. Experimentally measured IC 50 records were filtered, standardized, and deduplicated; where standardized continuous potency values were available, compounds were labeled using fixed potency thresholds after excluding a gray zone, while for curated datasets with discrete activity classes, intermediate entries were removed and binary labels were derived from the provided class annotations. Each molecule was encoded using a hybrid structural representation that concatenates ECFP4 fingerprints with MACCS keys, capturing complementary topological and predefined substructure information. We benchmarked RandomForest, ExtraTrees, LightGBM, a truncated-SVD logistic baseline (SVDLR), and equal-weight and OOF-weight-optimized probability ensembles under two held-out evaluation settings: a random stratified split and a Murcko scaffold-based split. Performance was assessed using ROC-AUC, PR-AUC, accuracy, balanced accuracy, F 1, MCC, Brier score, and expected calibration error. Across targets, the hybrid-fingerprint framework showed strong held-out discrimination for CYP19A1 and MMP-9, with more conservative but still informative performance for MMP-2 under both random and scaffold separation. These results support using the framework as an early-stage ranking and triage tool rather than as a direct proxy for therapeutic success. For CYP19A1, ligand-based potency prediction aligns with a well-established therapeutic class; for MMP-2/MMP-9, however, potency is not sufficient for translation due to historical broad-spectrum metalloproteinase binding and toxicity. Accordingly, we position our MMP models as potency-first triage tools, with shortlisted hits requiring selectivity counterscreens and liability-aware filtering prior to any therapeutic interpretation. • Curated ChEMBL IC 50, datasets for CYP19A1, MMP-2, and MMP-9 using stringent standardization and gray-zone exclusion to reduce label ambiguity. • Proposed a hybrid structural encoding by concatenating ECFP4 (2048-bit) and MACCS (167-bit) fingerprints to capture complementary substructure signals. • Benchmarked tree-based, linear, and ensemble models under two held-out generalization settings: a random stratified split and a Murcko scaffold-based split. • Reported ranking, threshold-dependent, and calibration metrics (ROC-AUC, PR-AUC, ACC, BACC, F 1, MCC, Brier score, and ECE) to provide a deployment-relevant assessment for virtual screening. • Adds assay-context summaries, endpoint-coverage reporting, calibration analysis, applicability-domain analysis, and liability-motif analyses to better contextualize potency predictions and assay heterogeneity.
Ahmad et al. (Sat,) studied this question.