Estimating causal treatment effects in observational studies is challenging, particularly when data exhibit hierarchical or clustered structures.Standard propensity score (PS) methods often fail to account for betweencluster variation, leading to biased estimates.Although trimmed IPW and overlap Weights (IPW-T and OW) reduce bias with parametric PS models, their performance with flexible nonparametric estimators has not been systematically examined.We provide the first comprehensive assessment of parametric and nonparametric PS models-including logistic regression with fixed and random effects, generalized boosted models (GBM), Bayesian additive regression trees (BART), and ensemble learners-paired with four weighting strategies: marginal IPW, clustered IPW, IPW-T, and OW.Monte Carlo simulations mimicking realistic clustered settings with nonlinear PS relationships and unmeasured cluster-level confounding were conducted.We considered both homogeneous treatment effects, where individuals share the same treatment effect, and heterogeneous treatment effects, where effects vary according to individual covariates.Under homogeneous effects, fixed-effects logistic regression and fixed-effects GBM with IPW-T or OW produced accurate estimates even with unmeasured confounding.Under heterogeneous effects, nonparametric models underestimated extreme PS values and yielded biased estimates, whereas fixed-effects logistic regression with IPW-T or OW remained more stable.IPW-T and OW consistently improved stability by reducing extreme weights but introduced residual bias by altering the target population.Overall, our study highlights the advantages and limitations of combining nonparametric PS estimation with IPW-T and OW, and offers guidance for causal inference in clustered observational data.
Choi et al. (Tue,) studied this question.