Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
May 6, 2026Scientific ReportsOpen Access

Measuring deep learning performance - an empirical study of performance distributions across architectures and tasks

View Full Paper
Ask AI
Bookmark
Share

Authors

KCKevin L. CoakleyOGOdd Erik Gundersen

Discussion

Loading...

Member takes

Overview

Analysis reveals how performance distributions differ across architectures and tasks in deep learning, indicating robustness and tail risk implications.

Key Points

  • This research aims to examine the impact of non-determinism on deep learning performance distributions across various architectures and tasks.
  • Conducted 186 experiments on different deep learning architectures for image classification and time series forecasting.
  • Executed each experiment 100 times with varying random seeds to create performance distributions.
  • Quantified robustness using metrics for spread, symmetry, and tail risk.
  • Performance distributions are often non-Gaussian, especially in time series forecasting.
  • Time series models exhibit significantly higher tail risk, with nearly three times more underperforming outliers compared to image classification models.
  • Mean performance does not reliably predict robustness, indicating the need for distributional analysis for model selection.

Cite This Study

Coakley et al. (2026) studied this question.

synapsesocial.com/papers/69fa98bd04f884e66b532802https://doi.org/10.1038/s41598-026-49656-z
View Full Paper
Ask AI
Bookmark
Share