Motor Third-Party Liability (MTPL) insurance is a challenging domain for predictive modeling due to extreme class imbalance, temporal drift, and strict interpretability and probability calibration requirements in regulated environments. This study proposes a time-aware hybrid risk modeling framework that integrates actuarially grounded statistical methods with modern machine learning techniques. Using a large-scale real-world MTPL dataset from 2016–2024, we combine GLM-based segmentation and contrast analysis with ensemble classifiers, including Random Forest, XGBoost, and Histogram-based Gradient Boosting (HistGB), under strict chronological validation. Models are trained on 2016–2021 data, calibrated on 2022, and evaluated on an out-of-time test set covering 2023–2024. Results show that temporal drift substantially reduces discrimination across all models (ROC AUC ≈ 0 . 62 –0.63), while meaningful differences emerge in calibration quality and screening behavior. HistGB provides the most well-calibrated probability estimates (Brier ≈ 0 . 026 ; ECE ≈ 0 . 04 ) and the highest recall among standalone classifiers under operationally constrained threshold selection. Building on these findings, a hybrid architecture combining One-Class Support Vector Machine (OCSVM) profiling with HistGB classification is introduced as a screening-oriented decision system. Rather than maximizing overall accuracy, the hybrid model integrates anomaly-style profiling with calibrated probability estimation to support early identification of high-risk policies. Interpretability is supported through GLM segmentation, Partial Dependence Plots, SHAP analysis, and surrogate decision trees. The proposed framework demonstrates how statistical transparency and machine-learning flexibility can be combined to enable explainable, operationally viable risk screening in regulated insurance environments. • Large-scale longitudinal MTPL dataset from a real-world insurance portfolio. • Time-aware evaluation framework aligned with actuarial practice and temporal risk drift. • Comparative analysis of statistical, machine learning, and hybrid screening-oriented models under severe class imbalance. • Well-calibrated HistGradientBoosting model supporting recall-focused operational risk screening. • Hybrid OCSVM-HistGB framework enabling interpretable and capacity-controlled claim risk identification.
Schmidt et al. (Wed,) studied this question.