PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 20, 20260 citationsOpen Access

A CNN2D-LSTM Framework for Rule-Based Pedestrian-Vehicle Risk Scenario Detection

OBOumaima BenkhaddaMMMeriem Mandar

Key Points

  • The aim is to develop a framework that automatically detects risky scenarios involving pedestrians and vehicles from video data.
  • Utilized a CNN2D for extracting visual features from video frames.
  • Employ LSTM to model temporal dependencies of the sequences.
  • Tested the approach using video data from the JAAD dataset.
  • Achieved an overall accuracy of 97% in risk detection.
  • Class-wise precision reached 99.5% for 'No Risk' class and 92.2% for 'Risk' class.

Abstract

In this work, we developed an approach for rule-based pedestrian-vehicle risk scenario detection from video sequences. The contributions lie in the classification of "risky" and "non-risky" situations automatically derived from behavioral and physical cues, which could improve accident prevention and intelligent driver-assistance systems. A two-dimensional convolutional neural network (CNN2D) is employed over the frames of the videos for visual features extraction, while the LSTM recurrent network models the temporal dynamics of the sequences. The data used in these experiments are sequences of video frames extracted from the JAAD dataset. Behavioral and physical variables include pedestrian crossing, gaze direction, vehicle action, proximity, which are used only for generating the risk labels. They are not provided as explicit input to the model. While these labels are heuristically generated and may not capture all possible risky scenarios, they provide a practical framework for model training and evaluation. The CNN2D extracts the spatial visual features from the frames, while the LSTM captures temporal dependencies, which permits the model to learn both in the spatial and temporal axes for the prediction of risk. Tests on the JAAD dataset composed of varied traffic conditions and images of pedestrian crossings report overall accuracy of 97% and class-wise precision 99.5% for the "No Risk" class and 92.2\% for the "Risk" class. These results confirm the effectiveness of the suggested model and demonstrate the usefulness of fusing visual and temporal information collectively for automatic risk detection in difficult traffic environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Benkhadda et al. (2026) studied this question.

synapsesocial.com/papers/6a0d5132f03e14405aa9da8dhttps://doi.org/10.19139/soic-2310-5070-3458
Ask AI
Helpful
Bookmark
Share
View Full Paper