KI for Disruptive Robotics

  • home_icon HOME.
  • Research
  • Research Institutes
  • KI for Disruptive Robotics
  • Highlights

Total 19

Making End-to-End Driving Robust to Camera Viewpoint Changes




Prof. Kuk-Jin Yoon’s group at KAIST has developed VR-Drive, a new end-to-end autonomous driving framework designed to handle changes in camera viewpoints across different vehicle platforms and sensor configurations. As end-to-end autonomous driving continues to gain attention for integrating perception, prediction, and planning within a unified architecture, one important challenge remains. In real-world deployment, camera positions often vary depending on the vehicle type and sensor setup, which can significantly degrade driving performance.

To address this problem, the KAIST team introduced a scalable framework that improves viewpoint generalization through 3D-aware view synthesis. VR-Drive jointly learns 3D scene reconstruction as an auxiliary task and uses the reconstructed scene to generate new training views that are relevant for driving decisions. This allows the model to better adapt to viewpoint changes without requiring additional annotations or expensive optimization during training.

At the core of VR-Drive is a feed-forward 3D Gaussian Splatting module that reconstructs scene geometry and appearance from sparse multi-view inputs. Unlike previous methods that depend on scene-specific optimization or computationally heavy rendering, VR-Drive performs view synthesis in a feed-forward manner, making it practical for online viewpoint augmentation. These synthesized views are then incorporated into the end-to-end perception and planning pipeline, helping the model learn to make reliable driving decisions even under unseen camera viewpoints.

To further improve consistency across viewpoints, the researchers introduced two complementary techniques. The first is a viewpoint-mixed memory bank, which enables temporal interaction between features extracted from different viewpoints during sequential training. The second is viewpoint-consistent distillation, which transfers knowledge from original-view features to synthesized-view features to reduce noise and encourage better alignment between representations. Together, these components help the model learn more stable and robust features across varying camera views.

Extensive experiments demonstrate that VR-Drive significantly improves planning performance under novel camera viewpoints compared to existing E2E-AD baselines. Notably, the model maintains strong performance even when evaluated on camera configurations that differ substantially from those seen during training. To facilitate systematic evaluation, the team also introduces a new benchmark designed to assess viewpoint generalization in autonomous driving scenarios.

This work, accepted to NeurIPS 2025, represents an important step toward practical deployment of end-to-end autonomous driving systems. By combining efficient 3D reconstruction, planning-aware view synthesis, and cross-view representation alignment within a unified framework, VR-Drive provides a scalable solution for handling real-world variability in sensor configurations and paves the way for more robust autonomous driving technologies.
2026 KI Newsletter


KAIST 291 Daehak-ro, Yuseong-gu, Daejeon (34141)
T : +82-42-350-2381~2384
F : +82-42-350-2080
Copyright (C) 2015. KAIST Institute