This thesis presents a framework that integrates anomaly detection into reinforcement learning (rl) systems operating in noisy and uncertain environments. the primary focus is on accurately identifying anomalous behavior by leveraging latent disagreement ensembles, which capture deviations in the model’s latent representations. the framework is structured around a mape-k loop (monitor, analyze, plan, execute, knowledge) that continuously assesses system performance and triggers adaptive responses when anomalies are detected. although the original dreamer v3 architecture is retained—with both a fast, cnn-based model and a more robust, masked vision transformer (vit) model—the vit with masking is now introduced primarily as an auxiliary mechanism for increased robustness rather than as the central focus. experimental results on atari games and simulated industrial environments demonstrate that incorporating an anomaly detection mechanism significantly improves the overall reliability and performance of rl agents.
Κωνσταντίνος Κ. Σπυριδόπουλος (Wed,) studied this question.