Vision–Language–Action (VLA) policies promise flexible long-horizon manipulation, but deployment under domain shift requires both reliable uncertainty estimates and a workable runtime-assurance policy. We study a model-agnostic uncertainty-calibrated safety-gating wrapper that estimates online failure risk and routes control among policy execution, pause-and-reobserve, and a fallback planner. Using a cleaned and consistently aggregated benchmark pipeline, we evaluate two long-horizon manipulation tasks in NVIDIA Isaac Sim 5.0 under lighting, texture, occlusion, sensor, and combined shifts. Relative to an ungated VLA baseline, calibrated gating improves mean shifted success from 57.5% to 77.2% and reduces aggregate expected calibration error from 0.303 to 0.100. The largest success gains occur under occlusion and combined shift, including improvements from 48.3% to 85.2% on the drawer task and from 59.4% to 87.8% on clutter sort. The results also expose a systems trade-off: an aggressive uncalibrated threshold baseline attains stronger raw success and collision metrics, but requires nearly twice as many interventions per shifted episode (21.6 vs. 11.5). The main contribution is, therefore, an empirical characterization of the reliability–intervention trade-off created by calibrated supervision, not a claim that the calibrated supervisor is universally the best terminal controller. We frame calibrated gating as a better-calibrated, lower-intervention supervisor that materially improves robustness relative to an ungated VLA while revealing the open problem of mapping calibrated risk into efficient intervention policies. Additional threshold-sensitivity, signal-diagnostic, overhead, and residual-failure analyses show that the selected operating point is meaningful but not universal: the calibrated risk threshold captures most shifted failures in retrospective logs, yet residual contacts still arise during pause and fallback states. These findings provide controlled simulation evidence for trustworthy VLA supervision under distribution shift and clarify the reliability–intervention frontier that future embodied-control systems must navigate.
Ghaleb et al. (Fri,) studied this question.