Evaluating the effectiveness and limitations of ChatGPT, Gemini, and DeepSeek in benchmarking the execution of deterministic eligibility rules for lung cancer screening from a structured patient list. 1,000 simulated patient cases (500 smokers and 500 non-smokers) were included. Three large language models (LLMs) were evaluated: ChatGPT-4.0, Gemini Advanced, and DeepSeek-V3. Various personal and smoking-associated characteristics, including age, gender, weight, height, smoking status (smoker versus non-smoker), packs per day, years smoked, pack years, and years since quitting smoking, were included in a file and presented to the LLMs. Eligibility criteria from three countries were included: the USA, South Korea, and Germany. The LLMs were evaluated for robustness, repeatability, factual correctness, and diagnostic accuracy after simple and advanced prompting. ChatGPT-4.0 achieved a specificity of 94.6% and 100% for lung cancer screening eligibility in the USA and a specificity of 97.3% and 100% for lung cancer screening eligibility in South Korea, after simple prompting. Gemini Advanced reached 100% accuracy after simple prompting. Advanced prompting led to accuracies of 100% using all three LLMs. DeepSeek failed to provide robust and repeatable results. Factual correctness was not offered for Germany by any of the LLMs. While ChatGPT-4.0 and Gemini Advanced can accurately execute deterministic screening rules from structured lists when provided with explicit instructions, their role in clinical practice remains supportive, as these results represent performance in an idealized and noise-free environment.
Jang et al. (2026) studied this question.