PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 10, 2026Scientific Reports1 citationsOpen Access

Automated phenotyping of ophthalmologic diseases from routine medical records using small language models and the human phenotype ontology (HPO)

View Full Paper
BTBinh Duong ThaiSASebastian ArensTRThomas Reinhard

Key Points

  • This research examines the effectiveness of the human phenotype ontology (HPO) in automated phenotyping within ophthalmology.
  • Developed an AI pipeline that integrates text segmentation and negation detection using a small language model (PHI-4).
  • Manually annotated 175 ophthalmic medical records with HPO terms to create a ground truth dataset.
  • Evaluated the AI pipeline's performance based on metrics like Jaccard similarity, precision, recall, and F1 score.
  • Identified a total of 342 HPO terms manually, with the AI pipeline retrieving 341 terms.
  • Achieved a median Jaccard similarity of 0.67, precision of 0.83, recall of 0.82, and F1 score of 0.80.
  • The AI pipeline effectively extracts standardized phenotypes, indicating potential for improved data management.

Abstract

Abstract Automated phenotyping in ophthalmology requires accurate standardization of clinical terms to facilitate interoperability and research. This study evaluates the suitability of the human phenotype ontology (HPO) for automated extraction of ophthalmic phenotypes from narrative documentation. We developed a locally operated AI pipeline combining text segmentation and negation detection based on a small language model (PHI-4) with a dense retrieval approach using an augmented multilingual HPO catalog. Synonyms were incorporated into the HPO during training on anonymized consecutive physician letters. To validate the pipeline, 175 anterior segment and fundus descriptions from randomly picked medical records were manually annotated with HPO terms as ground truth. Overall, 342 HPO terms were identified manually (on average 2.53 terms per document), with 341 retrieved by the pipeline (on average 2.52 terms per document). Performance metrics showed a median Jaccard similarity of 0.67, precision of 0.83, recall of 0.82, and F1 score of 0.80. These results demonstrate that our AI pipeline effectively extracts standardized HPO terms from free-text ophthalmic findings. Integration of this pipeline into clinical information systems may enhance data interoperability and reduce manual coding workload in ophthalmology practices in the future.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Thai et al. (2026) studied this question.

synapsesocial.com/papers/6a0020cec8f74e3340f9b9b5https://doi.org/10.1038/s41598-026-51512-z
Ask AI
Helpful
Bookmark
Share
View Full Paper