PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Programmatic Video Prediction Using Large Language Models

View Full Paper
JTJing TangKEK. V. EllisSLSuhas Lohit

Key Points

  • ProgGen demonstrates superior video frame prediction, outperforming existing techniques in multiple environments.
  • The method utilizes large language models to create human-interpretable states for predicting future video frames.
  • Empirical evaluations in PhyWorld and Cart Pole show significant enhancements in video generation tasks.
  • Counter-factual reasoning features enable further interpretability and application versatility in video-related tasks.

Abstract

The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applications such as video surveillance, robotics applications, autonomous driving, etc. this objective entails synthesizing plausible visual futures, given a few frames of a video to set the visual context. Towards this end, we propose ProgGen, which undertakes the task of video frame prediction by representing the dynamics of the video using a set of neuro-symbolic, human-interpretable set of states (one per frame) by leveraging the inductive biases of Large (Vision) Language Models (LLM/VLM). In particular, ProgGen utilizes LLM/VLM to synthesize programs: (i) to estimate the states of the video, given the visual context (i.e. the frames); (ii) to predict the states corresponding to future time steps by estimating the transition dynamics; (iii) to render the predicted states as visual RGB-frames. Empirical evaluations reveal that our proposed method outperforms competing techniques at the task of video frame prediction in two challenging environments: (i) PhyWorld (ii) Cart Pole. Additionally, ProgGen permits counter-factual reasoning and interpretable video generation attesting to its effectiveness and generalizability for video generation tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tang et al. (2025) studied this question.

synapsesocial.com/papers/68f5a78aab63786de5b46136https://doi.org/10.48550/arxiv.2505.14948
Ask AI
Helpful
Bookmark
Share
View Full Paper