PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 29, 2026Journal of Clinical Oncology0 citations

Identifying cancer cachexia using clinical reasoning artificial intelligence (AI).

View Full Paper
VGVibhor GuptaJZJingqing ZhangMYMichael Steven Yule

Key Result

A Clinical Reasoning AI model accurately identified documented cancer cachexia with 1.00 sensitivity and 0.982 F1 score, outperforming standard ICD coding and simple NLP methods.

Key Points

  • This study aims to determine whether Clinical Reasoning AI can accurately identify cancer cachexia using established guidelines.
  • Clinical Reasoning AI assessed cancer cachexia using structured and unstructured data from electronic health records.
  • The model analyzes patient information based on four cachexia guidelines through retrieval-augmented generation and natural language processing.
  • Accuracy was measured against a gold standard of 50 patients diagnosed by two independent clinicians.
  • The Clinical Reasoning AI showed high precision (0.762), sensitivity (1.00), and specificity (0.853) in identifying documented cachexia.
  • It outperformed baseline methods, particularly for patients meeting at least one guideline (Group 2).
  • ICD coding and simple NLP methods failed to detect patients classified in Group 2 due to their limited approaches.

Study Design

Type

Cross-Sectional (n=50)

Structured PICO

Does a guideline-driven Clinical Reasoning AI model improve the identification of cancer cachexia compared to standard ICD coding, NLP, and vanilla LLMs?

P
Population
50 patients evaluated for cancer cachexia, classified by two independent clinicians as a gold standard.
I
Intervention
Clinical Reasoning AI model employing retrieval-augmented generation and agentic techniques with LLMs, configured on four cachexia guidelines (Fearon, Evans, GLIM, mGPS) using structured and unstructured EHR data.
C
Comparator
ICD coding (ICD-9 799.4; ICD-10 C80.9), simple NLP using keywords, and Vanilla OpenAI GPT-o4-mini (without explicit guideline instructions).
O
Outcome
Accuracy (precision, sensitivity, F1 score, specificity) of classifying patients into three cachexia groups (Group 1: no features, Group 2: meets ≥1 guideline, Group 3: documented cancer cachexia).

A guideline-driven Clinical Reasoning AI model accurately identifies cancer cachexia from complex EHR data, outperforming standard ICD coding, keyword NLP, and vanilla LLMs.

Abstract

1613 Background: Cancer cachexia, a severe wasting syndrome, affects up to 87% of cancer patients 1 and contributes to ≥20% of cancer deaths 2 . Despite this burden, ~85% of healthcare providers are unfamiliar with diagnostic criteria 3 , leaving patients often undiagnosed. This study evaluated whether a Clinical Reasoning AI model configured on four cachexia guidelines (Fearon 4 , Evans 5 , GLIM 6 and mGPS 7 criteria) and designed to emulate clinician chart review, could accurately identify cancer cachexia. Methods: Using structured (vitals, labs) and unstructured (clinical notes) electronic health record data the Clinical Reasoning AI employed retrieval-augmented generation 8 and agentic techniques with large language models (LLMs) to: (1) retrieve clinically relevant information based on the clinical guidelines via natural language processing (NLP) and embedding, and (2) apply guideline-based reasoning to classify patients into Group 1 (no cachexia features), Group 2 (meets ≥1 guideline regardless of documented diagnosis), or Group 3 (documented cancer cachexia). Accuracy was assessed against a gold standard of 50 patients classified by two independent clinicians, with conflicts resolved by a third. Performance was compared against baseline methods: ICD coding (ICD-9 799.4; ICD-10 C80.9), NLP using keywords like cancer or cachexia, and LLMs (Vanilla OpenAI GPT-o4-mini) prompted without explicit instruction to use guidelines. Results: The Clinical Reasoning AI demonstrated high precision, sensitivity, and specificity, outperforming baselines, particularly for Group 2 patients. ICD and NLP methods could not detect Group 2, due to code- and keyword-based constraints. Performance of LLMs without guidelines worsened with structured and unstructured data, indicating sensitivity to data overload and inability to prioritise clinically relevant information. Conclusions: These results show that embedding clinical guidelines within a structured retrieval-and-reasoning architecture enabled accurate identification of cancer cachexia, addressing a critical gap where standard coding and keyword methods fail. The guideline-driven approach maintained accurate clinical reasoning while processing complex EHR data. The Clinical Reasoning AI model is currently being deployed at the point of care to assess real-world implementation and impact on patient outcomes. Group 2 Group 3 AI Model Precision Sensitivity F1 Specificity Precision Sensitivity F1 Specificity ICD N/A N/A N/A N/A 0.759 0.815 0.786 0.696 Simple NLP N/A N/A N/A N/A 0.727 0.593 0.653 0.739 Vanilla OpenAI GPT-o4-mini (unstructured data only) 0.500 0.500 0.500 0.765 0.594 0.704 0.644 0.438 Vanilla OpenAI GPT-o4-mini (structured and unstructured data) 0.500 0.062 0.111 0.971 0.385 0.185 0.250 0.652 Clinical Reasoning AI with GPT-o4-mini 0.762 1.00 0.865 0.853 0.964 1.00 0.982 0.957

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gupta et al. (2026) conducted a cross-sectional in Cancer cachexia (n=50). Clinical Reasoning AI model vs. ICD coding, simple NLP, and vanilla LLMs was evaluated on Accuracy of identifying cancer cachexia (precision, sensitivity, F1, specificity). A Clinical Reasoning AI model accurately identified documented cancer cachexia with 1.00 sensitivity and 0.982 F1 score, outperforming standard ICD coding and simple NLP methods.

synapsesocial.com/papers/6a192f07fab5b468c441844fhttps://doi.org/10.1200/jco.2026.44.16_suppl.1613
Ask AI
Helpful
Bookmark
Share
View Full Paper