PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 15, 2026IEEE Journal of Biomedical and Health Informatics0 citations

Dense Retrieval for Electronic Health Record With Knowledge Injection and Synthetic Data

View Full Paper
ZZZhengyun ZhaoHYHuaiyuan YingSYSheng Yu

Key Points

  • The aim is to enhance electronic health record retrieval by addressing semantic gap challenges using dense retrieval methods.
  • Developed DR.EHR, a series of dense retrieval models tailored for electronic health records.
  • Implemented a two-stage training pipeline using MIMIC-IV discharge summaries for entity extraction and knowledge injection.
  • Trained models with 110 M and 7B parameters, evaluated against the CliniQ benchmark for performance assessment.
  • DR.EHR significantly outperformed existing dense retrieval models on the CliniQ benchmark, achieving state-of-the-art results.
  • Models excelled particularly in challenging semantic matches, such as implication and abbreviation.
  • Ablation studies confirmed the efficacy of pipeline components, and models demonstrated generalizability on EHR QA datasets.

Abstract

Electronic Health Records (EHRs) are pivotal in clinical practices, yet their retrieval remains a challenge mainly due to semantic gap issues. Recent advancements in dense retrieval offer promising solutions but existing models, both general-domain and biomedical-domain, fall short due to insufficient medical knowledge or mismatched training corpora. Previous EHR retrievers generally focus on a small set of queries and fail to generalize. This paper introduces DR.EHR, a series of dense retrieval models specifically tailored for EHR retrieval. We propose a two-stage training pipeline utilizing MIMIC-IV discharge summaries to address the need for extensive medical knowledge and large-scale training data. The first stage involves medical entity extraction and knowledge injection from a biomedical knowledge graph, while the second stage employs large language models to generate diverse training data. We train two variants of DR.EHR, with 110 M and 7B parameters, respectively. Evaluated on the CliniQ benchmark, our models significantly outperforms all existing dense retrievers, achieving state-of-the-art results. Detailed analyses confirm our models' superiority across various match and query types, particularly in challenging semantic matches such as implication and abbreviation. Ablation studies validate the effectiveness of each pipeline component, and supplementary experiments on EHR QA datasets demonstrate the models' generalizability on various EHR corpus and natural language questions including complex ones with multiple entities. This work significantly advances EHR retrieval, offering a robust solution for clinical applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhao et al. (2026) studied this question.

synapsesocial.com/papers/6a06b7a1e7dec685947aa58bhttps://doi.org/10.1109/jbhi.2026.3692546
Ask AI
Helpful
Bookmark
Share
View Full Paper