PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 15, 20260 citationsOpen Access

Mind the Ladder: A Benchmark for Level 1--3 Causal Reasoning in World Models

View Full Paper
DPDi Prodi Paolo

Key Points

  • To establish a benchmark for assessing levels of causal reasoning in latent world models using Pearl's Ladder of Causality.
  • Introduced Mind the Ladder as a diagnostic benchmark for latent world models.
  • Operated on three levels of causality (association, intervention, counterfactuals) in the latent space.
  • Validated on the Glitched Hue Two Room environment to test for causal disentanglement.
  • VoE surprise measure does not reliably indicate causal accuracy; many models show high surprise without passing counterfactual tests.
  • Models can have high surprise due to physical violations yet fail Level 3 counterfactual assessments.

Abstract

World models based on Joint-Embedding Predictive Architecture (JEPA) have demonstrated emergent physical understanding through Violation-of-Expectation (VoE) paradigms. However, the "surprise" metric used to evaluate these models conflates statistical novelty with genuine causal reasoning. This paper introduces Mind the Ladder, a diagnostic benchmark and metric suite for testing causal fidelity in latent world models. The framework operationalises Pearl's Ladder of Causality (Level 1: Association, Level 2: Intervention, Level 3: Counterfactuals) directly in the latent space of a trained world model, making it architecture-agnostic. Three novel metrics are proposed: AAP Surprise Ratio, Structural Invariance, and AAP Consistency Advantage all grounded in the LeWorldModel (LeWM) architecture. The benchmark is validated on the Glitched Hue Two Room environment, which tests causal disentanglement between spurious correlations and true causal mechanisms. Results show that VoE surprise alone is insufficient: a model can exhibit high surprise for physical violations while still failing Level 3 counterfactual tests.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Di Prodi Paolo (2026) studied this question.

synapsesocial.com/papers/6a06b940e7dec685947abdd7https://doi.org/10.5281/zenodo.20162155
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Crossing the Causal Ladder: Architectural Separation of World Model and Policy Achieves Level 2 Causal Inference and Level 3 Novel Intervention Generalisation2026
  2. 2The Illusion of Causality in LLMs: A Developmentally Grounded Analysis of Semantic Scaffolding and Benchmark–Capability Mismatches2026
  3. 3CausalBench: A Comprehensive Benchmark for Evaluating Causal Reasoning Capabilities of Large Language Models2024 · 44 citations
  4. 4RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation2025
  5. 5Causal Strengths and Leaky Beliefs: Interpreting LLM Reasoning via Noisy-OR Causal Bayes Nets2025