ABSTRACT In analytic survey inference, the attributes of units in a target or frame survey population are idealized as a sample from a superpopulation statistical model, and model parameters are estimated from survey data drawn from a probability sample of the frame population. The data structure in such survey inference consists of relevant attribute data together with survey weights associated with all survey respondents. These weights relate to the probability of inclusion of each unit within the respondent set, and they enable consistent estimation in large populations and samples of all frame‐population averages of functions of the unit attributes. However, even when this is assumed correct, model parameters such as those for within‐cluster dependence between survey attributes from distinct respondents may not be identifiable from survey data with weights. That is, even assuming the superpopulation model, with a parametric dependence structure for attributes within clusters, if sampled data are observed with precisely correct weights equal to the reciprocals of single‐inclusion probabilities or of conditional probabilities of inclusion given unit data, multiple distinct values of the parameters of the superpopulation model may yield the same likelihood for the data for some sample designs compatible with the weights. This article first describes the background and existing methods for the design‐based estimation of cluster‐level model parameters from survey data on a clustered superpopulation using single‐inclusion weights. Nonidentifiability results are presented rigorously as mathematical examples, proving that large‐sample consistent estimation of within‐cluster dependence parameters from survey data with single‐inclusion weights is not always possible. This article is categorized under: Statistical and Graphical Methods of Data Analysis > Sampling Algorithms and Computational Methods > Maximum Likelihood Methods
Eric Slud (Mon,) studied this question.