PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 2026npj Digital Medicine0 citationsOpen Access

Multimodal multi-instance learning for cardiopulmonary exercise testing performance prediction

ZHZhe HuangWPWeishen PanSAShudhanshu Alishetti

Key Result

A multimodal multi-instance learning model using echocardiography and health records predicted peak VO2 with an R² of 0.603 and identified high-risk heart failure patients with an AUROC of 0.849.

Key Points

  • The research aims to enhance peak oxygen consumption predictions in heart failure patients using accessible multimodal data.
  • Developed a multimodal multi-instance learning framework using transthoracic echocardiography and electronic health records.
  • Estimated peak oxygen consumption and identified high-risk patients through model training and validation.
  • Evaluated performance metrics, including R² and AUROC, against prior methodologies.
  • Achieved an R² of 0.603 in predicting peak oxygen consumption, outperforming previous models.
  • Secured an AUROC of 0.849 for identifying high-risk patients, indicating better classification accuracy.
  • External validation resulted in an R² of 0.541 and AUROC of 0.870, enhancing detection of patients needing advanced therapies.

Study Design

Type

Observational (n=1,127)

Multicenter

Yes

Structured PICO

Does a multimodal multi-instance learning model using TTE and EHR data improve the prediction of peak VO₂ and identification of high-risk heart failure patients compared to prior single-instance models?

P
Population
1,127 patients referred for cardiopulmonary exercise testing (CPET) across four New York-Presbyterian affiliated hospitals (1,000 in development cohort, 127 in external validation cohort).
I
Intervention
Multimodal multi-instance learning framework using transthoracic echocardiography (TTE) studies (2D cine loops, M-Mode, Spectral Doppler) and structured electronic health record (EHR) data.
C
Comparator
Prior single-instance learning approach and unimodal models.
O
Outcome
Peak oxygen consumption (peak VO₂) prediction (measured by R², RMSE, MAE) and high-risk patient identification defined as peak VO₂ ≤ 14 mL/kg/min (measured by AUROC, sensitivity, specificity, balanced accuracy, F1 score).surrogate

A multimodal multi-instance AI framework using routine echocardiography and EHR data can accurately predict CPET-derived peak VO₂ and identify high-risk heart failure patients, potentially expanding access to advanced risk stratification.

Main Result

Absolute Event Rate: 0.849% vs 0.836%

Limitations

  • Inherent limitations of saliency mapping techniques for interpretability
  • Lower predictive performance in older patients (age ≥ 60) likely due to underrepresentation in the development cohort
  • Retrospective design with potential confounding by indication
  • Dataset derived from four academic hospitals in the New York metropolitan area, limiting geographic and institutional generalizability
  • Relatively limited size of the external validation cohort
  • Echocardiograms and CPETs were closely timed but not simultaneous, allowing for potential clinical changes between studies
  • Model does not disentangle the different physiologic components of VO₂

Abstract

Heart failure (HF) is a progressive and fatal disease that affects nearly 7 million individuals in the United States, with prevalence expected to surpass 10 million by 2040. Cardiopulmonary exercise testing (CPET) represents the gold standard for assessing functional capacity and predicting survival outcomes among HF patients but its widespread use is limited by practical constraints. Here we introduce a multimodal multi-instance learning framework that predicts peak oxygen consumption (peak VO₂), a critical indicator from CPET, using the more accessible transthoracic echocardiography (TTE) studies and electronic health records (EHR). By modeling the cross-modal interactions and the multi-instance structure of TTE studies, our approach significantly improves predictive accuracy and generalization. The model achieves an R² of 0.603 in peak VO₂ prediction and AUROC of 0.849 in high-risk patient identification, surpassing prior work (R² = 0.529, AUROC = 0.836). On the external validation cohort, the model achieves an R² of 0.541 compared to 0.395 and an AUROC of 0.870 compared to 0.797 from previous work. The improved performance more accurately allows for identification of patients who may benefit from advanced heart failure therapies that otherwise may have been missed.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2026) conducted an observational in Heart failure (n=1,127). Multimodal multi-instance learning AI model (TTE and EHR) vs. Prior single-instance ensemble AI model was evaluated on High-risk patient identification (AUROC) and peak VO2 prediction (R²). A multimodal multi-instance learning model using echocardiography and health records predicted peak VO2 with an R² of 0.603 and identified high-risk heart failure patients with an AUROC of 0.849.

synapsesocial.com/papers/69a91da8d6127c7a504c0a11https://doi.org/10.1038/s41746-026-02493-w
Ask AI
Helpful
Bookmark
Share
View Full Paper