Skip to content
arXiv Robotics — research abstracts· Zihao Zhang, Haochen Tian, Tianyu Li, Changhui Jing, Jingliang He, Naisheng Ye, Ziyuan Pu, Zhenjie Yang·· 2 days agoEditorial score63

Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?

Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?

Summary

This research explores whether improved visual representations from foundation models consistently enhance end-to-end autonomous driving performance. The study introduces ViRA, a framework that aligns visual representations without changing planner architecture, and finds that auxiliary perception supervision reduces sensitivity to VFM target selection. The results suggest that integrating VFMs into autonomous driving requires careful consideration of target selection and planner supervision.

Source: arXiv Robotics — research abstracts · Read original article ↗

Loading article text…

What the source reports

Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.

  • paper: not_reported
  • code: not_reported
  • dataset: not_reported
  • weights: not_reported
Source excerpts and review record

Automatically extracted; no manual editorial approval recorded.

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).