PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 2026Frontiers in Digital Health5 citationsOpen Access

Ontology- and LLM-based data harmonization for federated learning in healthcare

View Full Paper
NKNatallia KokashLWLei WangTGTom Gillespie

Key Points

Key points are not available for this paper at this time.

Abstract

Introduction: Semantic heterogeneity across electronic health records (EHRs) limits scalable and privacy-preserving analytics in healthcare. While federated learning (FL) enables collaborative modeling without sharing raw data, it requires consistent, ontology-aligned representations. We present an ontology- and large language model (LLM)-based data harmonization approach to support secure, interoperable FL workflows. Methods: We propose a general two-step pipeline for converting or annotating clinical text into a predefined target ontology format. First, candidate concepts are retrieved from the target vocabulary using embedding-based similarity search or ontology cross-references. Second, an LLM acts as a semantic validator, accepting or rejecting candidates based on explicit equivalence or subsumption criteria. The approach is ontology-agnostic and configurable; mapping to MONDO and HPO is demonstrated as a real-world use case. Final accepted mappings were evaluated against independent human expert assessment. Results: Across two clinical datasets, expert-LLM agreement reached up to 92%, with overall performance ranging from 78% to 91% depending on candidate-generation strategy. Retrieval alone was insufficient for reliable mapping, whereas LLM-based validation substantially improved precision while complementary retrieval strategies improved recall. Discussion: The proposed pipeline transforms ontology-based harmonization from a manual expert task into a reusable and configurable workflow suitable for federated healthcare research. By combining high-recall retrieval with LLM-based semantic adjudication, the approach enables scalable, privacy-preserving conversion of heterogeneous clinical text into standardized representations across domains.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kokash et al. (2026) studied this question.

synapsesocial.com/papers/6a0711d685d51e7cc7583ce0https://doi.org/10.3389/fdgth.2026.1756555
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Advances and Open Problems in Federated Learning2020 · 5,451 citations
  2. 2Corrigendum to: Pharmacogenomic clinical decision support design and multi-site process outcomes analysis in the eMERGE Network2019 · 2 citations
  3. 3The Medical Dictionary for Regulatory Activities (MedDRA)1999 · 1,414 citations
  4. 4Unsupervised biomedical named entity recognition: Experiments with clinical and biological texts2013 · 248 citations
  5. 5AcetoBase Version 2: a database update and re-analysis of formyltetrahydrofolate synthetase amplicon sequencing data from anaerobic digesters2022 · 8 citations