Abstract We study the adaptive infinite-horizon discounted control problem for Piecewise Deterministic Markov Processes (PDMPs) using a Nonstationary Value Iteration (NVI) scheme. PDMPs, as introduced by Davis (Markov models and optimization, monographs on statistics and applied probability, Chapman and Hall, London, 1993) evolve deterministically between random jumps whose jump rate λ, transition measure Q, and cost C depend on an unknown parameter ^* β ∗. The proposed NVI algorithm recursively updates the value function using current parameter estimates, enabling online implementation. We show that, for any sequence of strongly consistent estimators \ ^*ₙ\ β n ∗ converging almost surely to ^* β ∗, the resulting policy is asymptotically optimal under the discounted criterion.
Costa et al. (2026) studied this question.