Abstract Background/Aims Connective tissue diseases (CTDs) like Sjögren’s disease (SjD), systemic lupus erythematosus (SLE), and systemic sclerosis (SSc) are complex autoimmune conditions with overlapping symptoms, significant morbidity, and reduced quality of life. Research is challenged by difficulties in accurately identifying CTD patients in electronic health records (EHRs) due to non-specific coding and variable data quality. Standard coding systems often fail to capture the diagnostic complexity, leading to misclassification. Algorithms used for identification show varying accuracy, with concerns regarding sensitivity and specificity. Despite this, few studies have evaluated the validity of such algorithms across healthcare systems. This systematic review (SR) aims to assess the accuracy of codes and algorithms used to identify CTDs in EHRs and administrative databases. Methods The SR was pre-registered on PROSPERO. Studies were included if they assessed the accuracy of codes and algorithms used to identify CTDs in EHRs/administrative databases against a reference standard (clinician diagnosis or classification criteria). We focused on selected CTDs from the ICD-10 M30-M36 classification: SLE, SjD, SSc, Polymyositis/Dermatomyositis, Mixed Connective Tissue Disease (MCTD), and Undifferentiated Connective Tissue Disease (UCTD). We searched MEDLINE/PubMed, EMBASE and CENTRAL. Records were managed through Covidence software and two review authors independently screened titles/abstracts and full-texts against eligibility criteria. Disagreements were resolved by a third author. Data extraction was performed by two independent reviewers, including study characteristics, algorithm definition, reference standard and accuracy measures (sensitivity, specificity, positive predictive values (PPV), negative predictive values (NPV)). Quality assessment was performed using the QUADAS-2 tool. Results From 3,166 records identified, 26 studies met the inclusion criteria. Most were retrospective cohort studies, primarily conducted in North America (n = 15) and Europe (n = 7), using administrative or EHR data. SLE was the most frequently validated condition (n = 18), followed by SSc (n = 6). Algorithms based only on diagnostic codes showed generally moderate to high specificity, while sensitivity varied widely (25-95%), depending on number of required codes. Adding laboratory data (e.g, ANA, dsDNA) or medication records improved PPVs to 70-95%. More complex rule-based or machine learning algorithms, including random forest and NLP models, achieved the highest performance overall, with PPV and specificity often exceeding 80%. Across diseases, algorithms requiring ≥2 diagnostic codes or specialist-confirmed diagnosis achieved better balance between sensitivity and PPV, highlighting the importance of combining structured codes with clinical and laboratory data for accurate case identification in EHRs. Conclusion The accuracy of algorithms for identifying CTDs in EHRs varies depending on data source and algorithm complexity. Disease code-based definitions demonstrate overall high specificity but variable sensitivity, while the addition of laboratory, medication, or specialist-confirmed diagnosis improved performance. Machine learning algorithms achieved the highest accuracy. Integrating multiple data domains appears to optimise the identification and validity of algorithms for identifying CTDs in large-scale administrative and EHR-based research. Disclosure C. Saka: None. T. Mani Prabu Kumar: None. D. Keane: None. P. Albert: None. B. Clyne: None. C. McCarhty: None. G. Tynan: None. N. Dunne: None. M. Flood: None. E. McCarthy: None. F. Moriarty: None.
Saka et al. (2026) studied this question.