Abstract Missing data are common in real-world studies, yet the underlying missingness structure is often unknown, bringing additional uncertainty before an appropriate inference method can be applied. In this paper, we systematically examine two sources of such uncertainty: (1) the missing mechanism(s) involved and (2) the specific functional form of the missing model within a given mechanism. Focusing particularly on settings involving missing not at random (MNAR) data, we propose a two-step mixture-structure-based method, including a model filtering pre-screening step. The tasks of handling both sources of uncertainty and conducting reliable inference are then unified within a single EM-based framework. The core of our method lies in constructing a two-layer postulated mixture, which can be viewed as deliberately introducing an overfitted mixture – thereby enhancing flexibility and robustness to uncertainty. We consider two general scenarios in which the true data follow either a mixture or a non-mixture structure, and establish an identification framework for continuous finite mixtures potentially subject to MNAR. Simulation studies and a real-data application to the Medical Expenditure Panel Survey (MEPS) are utilized to demonstrate the performance of our method.
Zhou et al. (2026) studied this question.