PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 3, 2026JMIR AI1 citationsOpen Access

Performance of Large Language Models vs Conventional Machine Learning for Predicting Clinical Outcomes With Limited Data: Comparative Study

EBErwan BiganSDStéphane Dufour

Key Points

  • This study investigates the advantages of large language models over conventional machine learning in predicting clinical outcomes, particularly with small datasets.
  • Compared two large language models with conventional machine learning algorithms.
  • Used three clinical datasets related to sepsis, gastric cancer, and acute leukemia.
  • Ensured datasets were published after LLM knowledge cutoff to avoid bias.
  • LLMs outperformed conventional ML algorithms for training sizes below 50 patients.
  • Key metrics included better receiver operating characteristic, F1-score, average precision, and balanced accuracy.
  • Contextual information was crucial for the performance advantage of LLMs.

Abstract

Background Machine learning (ML) can be used to predict clinical outcomes. Training predictive models typically requires data for hundreds or thousands of patients. Lowering this requirement to a few tens of patients would enable new applications in clinical trials (eg, optimizing the design of a phase III trial by training a predictive model on phase II data and applying it to synthetic phase III patients) or in clinical decision support systems (for rare diseases or narrow indications). Large language models (LLMs) have recently been shown to outperform conventional ML algorithms for predictions on tabular data when the train dataset is small. Objective This study aims to investigate the advantage of LLMs compared with conventional ML for predicting clinical outcomes by applying state-of-the-art models to recently published clinical datasets. Methods We compared 2 LLMs, 1 proprietary (from OpenAI) and 1 open source (from the Meta Llama family), with conventional ML classification algorithms to predict clinical outcomes using 3 recently published clinical datasets spanning distinct conditions (sepsis, gastric cancer, and acute leukemia). Datasets were chosen such that their publication date was after the LLM knowledge cutoff date to ensure that models were never exposed to these data during pretraining. Datasets were sampled to vary the training size. Results On average, the 2 tested LLMs perform better than conventional ML for training sizes below 50 patients, using the receiver operating characteristic area under the curve, the F1-score, the average precision, and the balanced accuracy as metrics. Contextual information was found to be key to this advantage. Conclusions These preliminary results may be further optimized. They already show the potential of LLM-based ML to enable new clinical use cases when data are available for only a few tens of patients.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bigan et al. (2026) studied this question.

synapsesocial.com/papers/69cf5ebc5a333a821460d43bhttps://doi.org/10.2196/83853
Ask AI
Helpful
Bookmark
Share
View Full Paper