Abstract Rationale Primary Ciliary Dyskinesia (PCD) is a rare genetic disorder characterized by a heterogeneous clinical presentation but currently lacks a diagnostic gold standard, making accurate diagnosis challenging. Current diagnosis relies on clinician interpretation of multiple complementary tests, as symptom overlap with other respiratory conditions complicates diagnoses based on clinical features alone. Given these diagnostic complexities and the increasing availability of large language models (LLMs), parents and families are more likely to seek information from LLMs. Therefore, it is essential to assess the accuracy and reliability of widely available LLMs for considering PCD in the diagnostic evaluation. Methods We evaluated four open-source LLMs for their ability to accurately recommend further diagnostic testing for PCD. Each model analyzed 28 de-identified initial pulmonology clinic visit notes from patients later confirmed with PCD. For each case, the model indicated whether further PCD evaluation was warranted (yes/no/uncertain) and provided a brief rationale with any suggested diagnostic tests. Because all cases had a confirmed PCD diagnosis, only sensitivity (the proportion of correctly identified cases warranting further testing) was assessed. Results Figure 1 shows the performance of 4 LLMs in terms of considering a diagnosis of PCD, and the ensemble approach was based on consensus between the 4 LLMs. Sensitivity between models varied significantly (p 0.001), and ranged from 0.48 - 1.00, with model Mistral-7B performing the best and GPT OSS 120B performing the worst. All models except GPT OSS 120B, as well as the ensemble majority vote, performed significantly better than chance (p 0.05). Interestingly, results indicate that smaller models may offer comparable utility to larger ones in screening for PCD. This could have important implications for resource-efficient clinical AI applications. Conclusions In using LLMs as a screening tool for PCD, high sensitivity is preferred as false negatives are more problematic than false positives. Although optimal sensitivity thresholds for LLMs are not well defined, standard PCD screening tests typically achieve sensitivity of 0.80-0.98. Thus, an LLM with a sensitivity higher than 0.80 has comparable usefulness. In our study, Llama 3 70B, Qwen 2.5 7B, Mistral-7B, and the ensemble all exceeded this threshold, supporting their use as effective screening tools. Further studies with larger and more diverse datasets are needed to validate LLM-based identification of PCD. Nevertheless, LLMs show promise as supportive tools for clinically complex cases where the diagnostic workup can be optimized based on available resources and further narrowing of their differential diagnoses. This abstract is funded by: None
Rajwal et al. (Fri,) studied this question.