PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 24, 2026Bioinformatics0 citationsOpen Access

BioTriplex : A Full-Text Annotated Corpus for Fine-Tuning Language Models in Gene-Disease Relation Extraction Tasks

View Full Paper
CCCharlotte CollinsPFPanagiotis FytasİKİlknur Karadeniz

Key Points

  • To develop and validate BioTriplex, a corpus for fine-tuning language models in extracting gene-disease relationships.
  • Created a dataset of 100 annotated biomedical articles focusing on gene-disease relations.
  • Annotated 21 subtypes of gene-disease relationships within the corpus.
  • Fine-tuned the LLaMA 3.1 8B model using BioTriplex for specific extraction tasks.
  • The fine-tuned model outperformed zero-shot and few-shot methods within its architectural framework.
  • Achieved greater granularity in classifying gene-disease relationship types than previous models.
  • Validated the effectiveness of BioTriplex as a comprehensive dataset for enhancing biomedical language models.

Abstract

Abstract Motivation Automatic information extraction from biomedical texts requires machine learning methodology that can recognise biomedical entities, characterise inter-entity relationships, and relate extracted information to specific research topics. Large language models (LLMs) excel in general tasks but perform less reliably in the biomedical domain, where texts are characterised by extensive technical terminology and semantic variations from general literature. There is an unmet need for annotated full-text datasets that can be used to fine-tune language models for significant biomedical applications. Here, we focus on extraction of the complex relationships between genes and diseases. Results We present BioTriplex, a corpus of 100 full-length biomedical research articles (comprising 604 subsection texts) manually annotated with disease names, genes, and 21 subtypes of disease-gene relationships. We employ BioTriplex to train the LLaMA 3.1 8B language model in gene-disease relation extraction. Our fine-tuned model outperforms zero-shot and few-shot approaches, both within the LLaMA 3.1 architecture and across the larger state-of-the-art LLMs GPT-4 and Claude Sonnet 3.7, and classifies gene-disease relation types with broader scope and greater granularity than previously described. These results validate BioTriplex as a useful full-text data resource and underscore the value of specialised datasets in fine-tuning language models for important biomedical tasks. Availability https://github.com/PanagiotisFytas/BioTriplex Supplementary information Supplementary data are available at Bioinformatics online.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Collins et al. (2026) studied this question.

synapsesocial.com/papers/69746149bb9d90c67120b36bhttps://doi.org/10.1093/bioinformatics/btag037
Ask AI
Helpful
Bookmark
Share
View Full Paper