PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 20250 citationsOpen Access

Two Causally Related Needles in a Video Haystack

View Full Paper
MLMiaoyu LiQCQin ChaoBLBoyang Li

Key Points

  • Models face challenges in extracting information from two locations in long videos, and performance declines with distance between those locations.
  • Causal2Needles benchmark results indicate that existing VLMs struggle with understanding causal relationships in human behaviors.
  • Experiment findings reveal that 2-needle questions effectively measure the limitations of current approaches in long-context video understanding.
  • The benchmark prevents textual bias by offering complementary question formats that test both visual and textual comprehension.

Abstract

Evaluating the video understanding capabilities of Video-Language Models (VLMs) remains a significant challenge. We propose a long-context video understanding benchmark, Causal2Needles, that assesses two crucial abilities insufficiently evaluated by existing benchmarks: (1) the ability to extract information from two separate locations in a long video and understand them jointly, and (2) the ability to model the world in terms of cause and effect in human behaviors. Specifically, Causal2Needles introduces 2-needle questions, which require extracting information from both the cause and effect human-behavior events in a long video and the associated narration text. To prevent textual bias, these questions comprise two complementary formats: one asking to identify the video clip containing the answer, and one asking for the textual description of an unrelated visual detail from that video clip. Our experiments reveal that models excelling in pre-existing benchmarks struggle with 2-needle visual grounding, and the model performance is negatively correlated with the distance between the two needles. These findings highlight critical limitations in current VLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68dc12c58a7d58c25ebb08cdhttps://doi.org/10.48550/arxiv.2505.19853
Ask AI
Helpful
Bookmark
Share
View Full Paper