Learning-based and adaptive proposal mechanisms can substantially accelerate Markov chain Monte Carlo (MCMC) sampling, but they also introduce a previously underappreciated source of irreproducibility: path-dependent adaptation of the proposal kernel. Here, we systematically analyze this effect using generator-matrix-based adaptive MCMC and demonstrate that naive on-the-fly learning leads to severe run-to-run variability and sampler freezing. We show that short training produces a degenerate frozen sampler and that apparent overadaptation effects reported in small-seed analyses are artifacts of insufficient replication: with 10 seeds across 7 training lengths, spread variability (SD) decreases monotonically with training duration. To remedy these pathologies, we propose a simple and theoretically clean strategy in which a generator matrix is pretrained on a representative structure and then shared as a fixed proposal kernel across runs and related structures. The shared-matrix approach reduces cross-run spread by 68% compared to independent learning, transfers effectively across the related DNA structures tested here (transfer gap <1 Å), and maintains comparable sampling quality. Notably, we demonstrate that ensemble averaging of generator matrices degrades rather than improves performance and that minimizing apparent reproducibility metrics can be misleading when the sampler itself is adaptive. Our results establish practical guidelines for training length selection and provide a framework for reproducible adaptive MCMC, with quantitative validation demonstrated for coarse-grained DNA systems.
Tanigawa et al. (Mon,) studied this question.