Background: Korea's National Health Insurance Claims Data (NHICD) are a critical resource for generating real-world evidence.However, discrepancies between diagnostic codes and actual clinical conditions may lead to misclassification bias.Therefore, validated operational definitions are essential to ensure internal validity.This study systematically reviewed validation studies of patient identification algorithms using the NHICD to guide future research.Methods: PubMed and KoreaMed databases were searched for studies published from the inception of these databases to December 31, 2025.We included studies that developed patient identification algorithms using the NHICD, validated them against reference standards, and reported quantitative validity measures.Two independent reviewers screened the literature, and 12 studies were included.Results: Most studies used medical records as the reference standard and reported sensitivity and positive predictive value (PPV).Algorithm validity varied according to disease type and design.For cancers and rare diseases, combining diagnostic codes with registration codes for rare and intractable diseases (RID) achieved high validity, with both sensitivity and PPV exceeding 98%.For chronic diseases, incorporating prescription data or visit frequency improved accuracy, although strict medication criteria reduced the sensitivity for mild cases.For musculoskeletal conditions, algorithms combining diagnosis and procedure or surgery codes indicated high validity for acute events but lower performance for outpatientmanaged conditions. Conclusion:No one-size-fits-all algorithm exists for patient identification in the NHICD.Researchers must develop operational definitions optimized for the specific clinical characteristics of the target disease.Strategic combinations of diagnostic codes with RID codes, prescriptions, or procedure codes are recommended to enhance the reliability and reproducibility of claims data-based research.
Lee et al. (Mon,) studied this question.