Background: Forecasting disease burden informs health system planning, yet most projections use single statistical methods. Modern deep learning architectures and pretrained foundation models have not been systematically compared for this task. Methods: We obtained GBD 2023 disability-adjusted life year (DALY) rates for 18 causes across five high-income countries (1990–2023; 90 time series). Nine forecasting models spanning four paradigms — statistical (auto-ARIMA, ETS, Prophet), machine learning (XGBoost), deep learning (temporal fusion transformer, N-BEATS, N-HiTS, PatchTST), and a pretrained foundation model (Chronos, zero-shot) — were evaluated against a naive persistence baseline. All 502 possible inverse-MAE-weighted ensemble combinations were tested. Models were retrained on 1990–2023 to project Australian burden through 2040 with sex and age stratification. Results: PatchTST achieved the lowest individual-model MAE (36.70 DALYs/100,000; skill score 0.144), followed by Chronos zero-shot (38.81; 0.095) and ETS (39.64; 0.075). A 5-model ensemble achieved the best overall MAE (33.01; skill 0.230). Chronos — requiring no training — produced the best-calibrated prediction intervals (62.9% empirical coverage for nominal 80%), far exceeding TFT (19.6%) and Prophet (23.1%). For Australia, anxiety (+111.5%) and dementia (+25.0%) showed the largest projected increases to 2040, while IHD (−41.9%) and stroke (−30.9%) showed the largest decreases. Female liver cancer DALYs were projected to increase 43.2% versus a 6.4% male decrease. Conclusions: PatchTST and pretrained foundation models outperformed all conventional methods for GBD forecasting, while multi-model ensembles provided the best overall accuracy. The zero-shot performance of Chronos demonstrates that pretrained time-series models transfer effectively to epidemiological domains without task-specific training. Code and data: https://doi.org/10.5281/zenodo.19528301
Hayden Farquhar (Sun,) studied this question.