Camera-based E2E (End-to-end) autonomous driving models map camera images directly to control commands. In onboard environments, real-time inference and driving performance are required. However, their performance depends on the resolution of the input image. High-resolution images preserve visual features of objects such as traffic lights and leading vehicles but increase computation and latency, while low-resolution images reduce latency but lose important visual information. This tradeoff between inference time and driving performance is an issue in onboard environments. This paper proposes a low- and high-resolution feature fusion module. The model extracts global features from low-resolution images and selectively fuses features from high-resolution regions of interest for traffic lights and leading vehicles. The proposed method was evaluated in the MORAI simulator and on the NVIDIA Orin AGX platform. The Driving Score, which evaluates driving performance, was 22.61% higher and the inference time was 35.91% faster than those of the model with high-resolution images. The proposed method alleviates the limitation of camera-based E2E caused by the tradeoff between inference time and driving performance in onboard environments.
Yeon et al. (2026) studied this question.