Sequence-specific protein-nucleic acid interactions underpin essential processes in gene regulation, yet generalizable computational methods for simultaneously predicting protein recognition sites and binding affinities with DNA and RNA remain limited, largely due to the sparsity of experimental binding data. To address this challenge, our group has developed biophysics-inspired, data-driven models for precise predictions of protein-nucleic acid interactions. By integrating higher order structural information with sequence features of protein-nucleic acid complexes into an optimized energy model, our model achieves state-of-the-art accuracy in predicting sequence-specific protein-DNA and protein-RNA binding affinities. Leveraging recent advances in AI-predicted complex structures, we further demonstrate the model’s applicability in cases where experimental structures are unavailable. Beyond quantitative affinity prediction, the model identifies favorable interaction motifs for given protein targets, facilitating the rational design of therapeutic nucleic acid aptamers. It also enables high-throughput genomic binding-site predictions for DNA-binding proteins and provides direct physicochemical interpretation of interactions between individual amino acids and nucleotides. Finally, we incorporate this data-driven model into a residue-resolution simulation framework that quantitatively captures absolute free energies of protein-nucleic acid binding. Together, our work establishes an integrative computational platform that reduces experimental costs in assessing nucleic acid recognition and enables mechanistic studies of various genetic and epigenetic processes.
Zhang et al. (Sun,) studied this question.