Vision-Language-Action in Robotics: A Survey of Datasets and Data Infrastructure
Overview
A 2026 TMLR survey paper reviews datasets, benchmarks, and data engines for Vision-Language-Action (VLA) in robotics, emphasizing challenges in data fidelity, scalability, and evaluation.
Generated from attributed reports · 33 minutes agoUpdated
Event evidence and corrections
0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.
No current evidence-backed claims. Missing information remains not reported.
Developments
- 2026-10-05 17:58 UTC · 1 reportsVision-Language-Action in Robotics: A Survey of Datasets, BeRobotics — Paper and dataset web discovery:Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
- 2026-10-05 17:58 UTC · 2 reportsVision-Language-Action in Robotics: A Survey of Datasets ...Robotics — Paper and dataset web discovery:Vision-Language-Action in Robotics: A Survey of Datasets and Data Infrastructure
Report timeline
Follow attributed reports and material updates.
- Robotics — Paper and dataset web discoveryVision-Language-Action in Robotics: Dataset Survey
Study reviews datasets, benchmarks, and data engines for Vision-Language-Action in robotics. Published in TMLR 2026.
- Robotics — Paper and dataset web discoverySignalVision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
This paper presents a systematic analysis of Vision-Language-Action (VLA) research, focusing on datasets, benchmarks, and data engines. It identifies key challenges in data fidelity, scalability, and evaluation, and proposes a structured approach to address these issues. The authors release an open-source repository to support the community.
- Robotics — Paper and dataset web discoveryVision-Language-Action in Robotics: A Survey of Datasets and Data Infrastructure
This survey paper examines the data infrastructure challenges in Vision-Language-Action (VLA) models for robotics. It categorizes datasets by embodiment diversity, modality composition, and action space formulation, identifies limitations in simulation-based and video-reconstruction paradigms, and outlines four open challenges: representation alignment, multimodal supervision, reasoning assessment, and scalable data generation.
Event attention history
There is not enough continuous observation data to show a trend.