This study evaluates the out-of-sample forecasting performance of realized volatility models using five-minute intraday data for the S&P 500 index and its constituent stocks over the period from 1997 to 2013. A structured econometric benchmark, the Heterogeneous Autoregressive model with realized quarticity (HARQ), which explicitly accounts for measurement error in realized volatility, is compared with machine learning and deep learning models that rely on the same information set but learn predictive patterns in a data-driven manner. Forecast evaluation is conducted under consistent training and testing schemes, and deep learning models are estimated over 40 repeated runs to account for stochastic optimization effects. Robustness analyses are performed across alternative window specifications and extended forecast horizons. The empirical results show that HARQ delivers strong and robust average forecasting performance across the full sample and most individual stocks. Deep learning models exhibit competitive forecasting performance under certain market regimes and at longer forecast horizons. Overall, the findings suggest that econometric and data-driven approaches offer complementary strengths in forecasting realized volatility.
Sanghyeon Kim (Thu,) studied this question.