Key points are not available for this paper at this time.
In this issue of PNAS, Gao et al. (1) probe the limits of Bayesian phylodynamic inference, a statistical framework that has revolutionized the study of pathogen evolution and epidemic spread.By exhaustively analyzing 15 landmark viral datasets with billions of Markov chain Monte Carlo (MCMC) iterations, they reveal that the "tree landscape" explored by standard inference algorithms is frequently rugged, with multiple peaks separated by deep valleys that chains rarely cross.Most strikingly, they show that just a handful of data points often drive these sampling failures, and that the consequences for biological conclusions can be substantial: Independent MCMC chains analyzing HIV subtype B sequences yielded divergence dates diering by decades, while the Lassa virus dataset showed most chains trapped in two distinct regions of tree space, half on one side and half on the other, with only two chains of the ten ever crossing between peaks.The same Lassa dataset, along with the Rabies dataset, mixed well within individual chains even though the replicate chains did not converge, suggesting that sampling pathologies may hide within single runs or insuicient numbers of independent replicates.Their findings force us to confront an uncomfortable reality: Even with extraordinary computational eort, the limits of current implementations of this powerful statistical framework demand attention.
Drummond et al. (Mon,) studied this question.