Key points are not available for this paper at this time.
Drug-target interaction (DTI) prediction is critical for candidate compound screening and elucidation of mechanisms of action in drug discovery and repurposing. However, existing methods often rely on unimodal representations, global fusion, or shallow cross-modal fusion, making it difficult to adequately model the heterogeneous and fine-grained dependencies between drug structures and protein sequences. To address this issue, we propose CM-MTL-DTI, a DTI-oriented collaborative alignment framework, rather than a simple combination of auxiliary modules. The framework employs two independent one-dimensional convolutional neural networks as main-task encoders to extract backbone sequential representations from drug SMILES sequences and protein sequences, respectively, and introduces a GIN-based graph encoder to provide a complementary structural perspective for the drug modality. On this basis, we design an asymmetric bidirectional cross-modal attention mechanism to explicitly model direction-sensitive dependencies between drug substructures and protein residues. Meanwhile, three collaborative objectives─cross-modal masked reconstruction (XMR), graph-sequence consistency learning (GSC), and supervised contrastive learning (SupCon)─are introduced to achieve local semantic recovery, multiview semantic alignment within the drug modality, and discriminative enhancement of interaction representations, respectively. Experimental results on three benchmark data sets show that CM-MTL-DTI delivers stable and competitive performance under both standard and challenging settings, validating the effectiveness of the proposed DTI-oriented collaborative design.
Zhao et al. (Mon,) studied this question.