Standard parametric regression models for interval-censored time-to-event data frequently depend on a restrictive assumption of population homogeneity. When applied to complex, heterogeneous populations containing unobserved latent subgroups, these traditional models invariably estimate an averaged hazard function. Consequently, this homogenization systematically biases survival projections, overestimating survival times for high-risk participants while yielding overly pessimistic prognoses for low-risk individuals. To address this fundamental limitation and the known instability of traditional iterative mixture models, we introduce the prior-stratified stochastic multiple imputation algorithm, designed to disentangle latent heterogeneity and robustly impute continuous event times. Because estimating expectation-maximization mixtures from sparse interval-censored data often results in component collapse, prior-stratified stochastic multiple imputation structurally bypasses the fragile iterative expectation-maximization loop. Instead, it utilizes a one-time static Bayesian risk stratification, seeded by clinical priors, to anchor the likelihood space. This explicitly decouples the mixture into independent, strictly convex Weibull regressions, guaranteeing stable convergence even under extreme missingness. To rigorously quantify estimation uncertainty, the algorithm subsequently executes multiple repeated stochastic draws from the inferred participant-specific truncated distributions. Simulation studies demonstrate that our algorithm successfully prevents component collapse under extreme censoring (more than 40%) and structural misspecification, maintaining high phenotypic identifiability where standard unanchored mixtures fail entirely. Empirical validation utilizing a semi-synthetic primary biliary cholangitis cohort, the signal tandmobiel dental emergence dataset, and the highly sparse Finkelstein breast cancer dataset confirms the algorithm’s robust capacity to autonomously recover latent risk architectures and track non-parametric (Turnbull) ground truths without supervised labeling. Our algorithm offers a rigorously quantified methodological bridge, converting complex, heterogeneous interval-censored observations into complete datasets, thereby unlocking conventional survival analysis toolkits while safely preserving biological dimorphism.
Kapoor et al. (2026) studied this question.