Background Cyclists are among the most vulnerable road users in urban traffic environments. For autonomous vehicles to interact safely and effectively with cyclists, perception systems must go beyond detection and segmentation to include an explicit understanding of cyclist orientation. However, most existing cyclist datasets lack synchronized inertial metadata describing body orientation, limiting their use in multimodal and orientation-aware perception studies. Methods This Data Note presents a multimodal visual–inertial dataset acquired using the CYCLIST+IMU framework, which synchronizes monocular RGB images captured from a vehicle-mounted camera with inertial measurements recorded by bicycle-mounted and vehicle-mounted inertial measurement units. Data were collected during multiple real-world urban acquisition sessions, resulting in 3,606 RGB images, each temporally aligned with inertial measurements, including cyclist orientation angles (yaw and roll). From these acquisitions, cyclist-centered image crops were generated and manually annotated, resulting in polygon-based semantic segmentation labels, region-of-interest detection files, and relative depth maps estimated from the RGB images. To improve angular coverage, a targeted data augmentation strategy based on horizontal image flipping was applied to underrepresented orientation ranges, resulting in the generation of 718 additional samples. The final dataset comprises 4,324 synchronized multimodal samples organized in a hierarchical directory structure that preserves one-to-one correspondence across all data modalities. Conclusions The CYCLIST+IMU dataset provides synchronized RGB image crops, inertial orientation metadata, semantic segmentation annotations, relative depth maps, and detection files for 4,324 cyclist instances captured under real urban traffic conditions. By explicitly integrating visual and inertial data with precise temporal alignment and detailed documentation, this dataset enables reproducible research on cyclist orientation estimation, semantic segmentation, and multimodal sensor fusion for intelligent transportation systems.
Gómez-Meneses et al. (Wed,) studied this question.