This repository contains the essential alignment database and dependency submodules required for the operation of Rider, a computational tool designed for RNA viral protein identification and analysis. This collection serves as a comprehensive resource for validating structural similarities between Rider-predicted candidates and known reference RNA viral proteins, specifically focusing on RNA-dependent RNA polymerase (RdRp). Related Publication:The preprint describing the Rider methodology and this dataset is available on bioRxiv: Title: Expanding the RNA Virus Universe by Scalable Structure-Guided Discovery (2025-11-29) Link: https://doi.org/10.1101/2025.11.24.690314 Key Components: RSDB (Rider Broad Structure Library):Contains 217,762 predicted protein structures in .pdb format. Sourced from multiple public datasets, this broad library includes a wide array of RNA virus-encoded proteins beyond canonical RdRps (e.g., helicases, proteases, and capsid proteins). It is designed to enable comprehensive structural comparisons and the detection of diverse viral fragments or mobile genetic elements. RDSDB (Rider RdRp-domain-specific Database):Contains 189,694 refined, domain-specific RdRp structures in .pdb format (modeled using ESMFold v1). To address the specificity challenges posed by multi-domain polyproteins, this database was constructed by rigorously scanning and extracting only the RdRp catalytic domains using HMM profiles. Extraneous non-RdRp regions were removed, making this dataset highly optimized for detecting fragmented viral sequences. RDSDB30 (Non-redundant RdRp Database):Contains 9,735 representative RdRp domain structures in .pdb format. This highly curated dataset was generated by clustering the RDSDB structures using Foldseek at a 30% sequence identity threshold. This domain-centric, redundancy-reduced database serves as a high-density, computationally efficient target for rapid structural alignments, particularly for short or fragmented metatranscriptomic contigs. Integrated Submodules: ESM2 (35M parameters): Pre-trained weights utilized for protein sequence tokenization and embedding generation. ESMFold (v1 model): Model weights for high-accuracy protein structure prediction. Foldseek: Binary/Executable components for fast and sensitive structural alignment. Usage Note: This version is released primarily to support the peer review process of the Rider manuscript. The repository will be updated with refined datasets and documentation upon the official publication of Rider.
Gaoyang Luo (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: