Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?
Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?
This research explores whether improved visual representations from foundation models consistently enhance end-to-end autonomous driving performance. The study introduces ViRA, a framework that aligns visual representations without changing planner architecture, and finds that auxiliary perception supervision reduces sensitivity to VFM target selection. The results suggest that integrating VFMs into autonomous driving requires careful consideration of target selection and planner supervision.
Source: arXiv Robotics — research abstracts · Read original article ↗
Loading article text…
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
- paper: not_reported
- code: not_reported
- dataset: not_reported
- weights: not_reported
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
Source:arXiv Robotics — research abstracts · arxiv.org