Introduction: Long non-coding RNAs (lncRNAs) are a major class of non-coding RNAs (ncRNAs) longer than 200 nucleotides. They play key roles in plant embryogenesis, root development, reproduction, and gene silencing. Accurate identification of lncRNAs in crop species is crucial for understanding their biological functions. However, most existing computational tools are designed for human and animal lncRNAs, limiting their applicability to crop genomes. Therefore, this study aims to develop a crop-specific computational tool for the accurate classification of lncRNAs and coding RNAs in crop species, addressing the limitations of current approaches. Methods: An XGBoost classifier was trained to distinguish lncRNA and coding RNA (cRNA) sequences using sequence-intrinsic features derived from five crop species: wheat, sorghum, rice, soybean, and maize. Model performance was evaluated against benchmark tools, CPC2 and PLEKv2. Results: The trained XGBoost classifier achieved an accuracy of 95.30%, precision of 93.90%, recall of 98.40%, F1-score of 96.10%, and an area under the ROC curve (AUC-ROC) of 99.40%, outperforming existing tools. These results demonstrate the model’s reliability in distinguishing lncRNAs from coding RNAs. Discussion: The trained XGBoost classifier was deployed as CLnc-Pred, a web-based application that allows users to input or upload FASTA sequences for lncRNA prediction. This framework enables efficient and accurate identification of lncRNAs in crop species. Conclusion: CLnc-Pred enhances accessibility and accuracy in crop lncRNA research and supports downstream functional and regulatory analyses. Future work will focus on expanding datasets, incorporating additional plant species, and extending the framework to support multi-class classification of diverse ncRNA types.
Choubisa et al. (2026) studied this question.