Structure-based drug discovery (SBDD) seeks ligands that bind to protein targets, but current workflows remain slow, costly, and limited by scarce high-quality data. While deep generative models show promise in designing ligands from protein structures, their performance is constrained by data scarcity and a lack of physical grounding. We introduce a physics-guided active learning framework that efficiently explores chemical space and generates biophysically realistic ligands with minimal data requirements. Our model extends conditional, variational, and autoencoder-GAN architecture to generate 3D ligands conditioned solely on receptor pocket geometry. Instead of relying on large static data sets, we implement an iterative active learning loop; in each round, the top 10% of generated ligands are selected using a composite physics-based score that integrates binding affinity (GNINA) and solvation energy (Amber), and then added back into training. This feedback mechanism refines the model toward physically viable and synthesizable candidates. Applied to the refined PDBBind v2019 data set, our framework achieves stronger binding across diverse targets, with median affinities of −9.62 kcal/mol versus −9.02 kcal/mol for references, and consistent gains even lower ranked ligands per target. We also observe an average synthesizability score (ASKCOS) improvement of 0.12e 4 , reflecting shorter and more feasible synthetic routes. On the BRD4 benchmark, the framework correctly ranks a known non-binder the lowest, validating the robustness of the scoring pipeline. Additionally, generated ligands exhibit more drug-like properties, including reduced molecular weight, improved lipophilicity, and lower polarity, factors that collectively favor bioavailability. Our physics-guided active learning framework overcomes data limitations in SBDD and enables biophysically informed exploration of chemical space. By grounding ligand design in binding energetics, solvation, and drug-likeness, it offers a practical route to molecules that are both computationally efficient and experimentally relevant.
Dhiman et al. (Sun,) studied this question.