DATA & TRAINING
Robotics datasets & training data
Human demonstrations, robot trajectories, synthetic data, teleoperation and annotation. Read original evidence before comparing scale or availability.
Origin
Human egocentric data, robot trajectories and synthetic data remain distinct.
Scale
Unique elapsed hours, summed sensor-hours and trajectory counts are different quantities.
Access
Availability and licensing require an explicit source. Missing information remains not reported.
Data and training coverage
All eligible data coverage appears here, including articles below the Signal selection threshold. Up to 50 latest records.
ETH Robotic Systems Lab — X
We are releasing EgoHTR, a dataset with both human motions and terrain references, accepted @corl_conf. 📖 Paper: https://lnkd.in/eeepUefs 🌐 Project Page: https://egohtr.github.io • 55 scene-aligned sequences • 150k+ frames • rough-terrain environments • multi-modal 3D scene
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Not reported
- Control:
- Not reported
- Data origin:
- Not reported
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “We are releasing EgoHTR, a dataset with both human motions and terrain references, accepted @corl_conf.”
Source E1
Original sourceNVIDIA — Robotics
NVIDIA has released an open-source, GPU-accelerated Medical Physics Simulation framework as part of its Isaac for Healthcare platform. This tool enables developers to model anatomy-device interactions, generate complex scenarios, and train robot policies in simulation, significantly reducing development time and improving regulatory readiness.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Not reported
- Control:
- Not reported
- Data origin:
- Not reported
Reported quantities Scroll across to read all columns.| Metric | Value / unit | Basis / context | Evidence |
|---|
| training environments | 8,192 robots | Reported trials Source wording: “benchmarks show 8,192 robot-training environments running in parallel” | Source E1 |
| training time | 5 hours | Unique elapsed hours Source wording: “training from over five hours to under two minutes” | Source E1 |
| training time | 2 hours | Unique elapsed hours Source wording: “training from over five hours to under two minutes” | Source E1 |
| clinical data | 500 hours | Summed sensor-hours Source wording: “nearly 500 hours of anonymized clinical data” | Source E2 |
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “benchmarks show 8,192 robot-training environments running in parallel with GPU-native simulation cut training from over five hours to under two minutes”
Source E1
Original source quotation: “CMR contributed nearly 500 hours of anonymized clinical data from its Versius Surgical Robotic System to the Open-H Embodiment open dataset”
Source E2
Original sourceMIT — Robotics
MIT researchers have developed FloatForm, a system of small robotic boats that can autonomously assemble into floating structures, disassemble, and reconfigure. Inspired by ant rafts, the system uses decentralized control and modular components to create dynamic, programmable water-based infrastructure. The work has potential applications in urban planning, emergency response, and adaptive public spaces.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Controlled physical settingSource E1
Reported quantities Scroll across to read all columns.| Metric | Value / unit | Basis / context | Evidence |
|---|
| assembly time | 4 hours | Unique elapsed hours experiments at MIT Source wording: “each run taking four to eight minutes” | Source E1 |
| success rate | 90 percent | Reported trials 10 trials Source wording: “completed its missions without human intervention 90 percent of the time with four robots” | Source E1 |
| success rate | 70 percent | Reported trials 10 trials Source wording: “completed its missions without human intervention 70 percent of the time with eight robots” | Source E1 |
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “In experiments at MIT, a fleet of eight robots repeatedly gathered from random positions into a target shape, latched into a rigid structure, broke apart on command, reassembled into a new configuration, and then drove across the pool as a single vessel, with each run taking four to eight minutes.”
Source E1
Original sourceUnitree Robotics — X
Unitree says it supports BitRobot in open-sourcing HIW-500, a humanoid teleoperation dataset collected across 12 real homes in Southeast Asia. The quoted announcement reports 500+ hours, 23K+ episodes, 10+ TB and more than 10 household tasks. The post does not specify the time basis or license, or link to a downloadable artifact.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Not reported
- Data origin:
- Not reported
Reported quantities Scroll across to read all columns.| Metric | Value / unit | Basis / context | Evidence |
|---|
| hours | 500 hours | Basis not reported Source reports 500+ hours; unique elapsed versus summed sensor-hours is not reported. Source wording: “500+ hrs” | Source E2 |
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “the largest open-source humanoid teleop dataset collected in real homes”
Source E1
Original source quotation: “500+ hrs
> 23K+ episodes”
Source E2
Original sourceHugging Face — Robotics
This tutorial explores the challenges of deploying VLA models on embedded robotic systems, including dataset recording best practices, fine-tuning techniques for ACT and SmolVLA, and real-time performance optimization using the NXP i.MX 95 SoC. It emphasizes asynchronous inference and hardware-aware scheduling to improve control and reduce latency.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Controlled physical settingSource E1
- Control:
- Not reported
Reported quantities Scroll across to read all columns.| Metric | Value / unit | Basis / context | Evidence |
|---|
| dataset episodes | 120 trajectories | Trajectory count Source wording: “Dataset: 120 episodes: 10 clusters x (10 different tea bag starting positions + 2 recovery episodes)” | Source E1 |
| inference latency | 0.32 hours | Unique elapsed hours Source wording: “i.MX 95 ACT Optimized 0.32 s” | Source E1 |
| inference latency | 2.86 hours | Unique elapsed hours Source wording: “i.MX 95 ACT ONNX FP32 2.86 s” | Source E1 |
| inference latency | 6.15 hours | Unique elapsed hours Source wording: “we have already established a baseline and measured an optimized on-board inference latency of 6.15 s” | Source E1 |
| test set accuracy | 1 percent | Summed sensor-hours Source wording: “Test Set (20) Accuracy: 1.00” | Source E1 |
| validation set accuracy | 0.9 percent | Summed sensor-hours Source wording: “Validation Set (10) Accuracy: 0.90” | Source E1 |
| global accuracy | 0.96 percent | Summed sensor-hours Source wording: “Global Accuracy (30): 0.96” | Source E1 |
| test set accuracy | 0.5 percent | Summed sensor-hours Source wording: “i.MX 95 SmolVLA ONNX FP32 29.1 s 0.50” | Source E1 |
| validation set accuracy | 0.4 percent | Summed sensor-hours Source wording: “i.MX 95 SmolVLA ONNX FP32 29.1 s 0.50 0.40” | Source E1 |
| global accuracy | 0.47 percent | Summed sensor-hours Source wording: “i.MX 95 SmolVLA ONNX FP32 29.1 s 0.50 0.40 0.47” | Source E1 |
- dataset: Not reported Source E1
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “Authors : Enzo Ruedas , Tess Boivin Recent advances in Large Language Models have enabled the transition from text-only reasoning to multimodal systems . First, with the integration of visual perception in Vision–Language Models (VLMs) , and more recently with the generation of robot actions in Vision–Language–Action (VLA) models . Deploying these models on embedded robotic platforms remains a challenge due to tight constraints in terms of compute, memory, and power, as well as real-time control requirements.”
Source E1
Original sourceHugging Face — Robotics
Hugging Face announces LeRobot v0.4.0, a major upgrade for open-source robotics with Dataset v3.0, new VLA models like PI0.5 and GR00T N1.5, and a plugin system for hardware integration. The release also adds support for LIBERO and Meta-World simulations, multi-GPU training, and a new Hugging Face Robot Learning Course.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Control:
- Not reported
Reported quantities Scroll across to read all columns.| Metric | Value / unit | Basis / context | Evidence |
|---|
| dataset size | 400 other | Unique elapsed hours Source wording: “datasets at the OXE-level (> 400GB)” | Source E1 |
| training time reduction | 2 other | Basis not reported Source wording: “cutting it in half with 2 GPUs” | Source E1 |
| training time reduction | 3 other | Basis not reported Source wording: “down to a third with 3 GPUs” | Source E1 |
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “datasets at the OXE-level (> 400GB)”
Source E1
Original source quotation: “LIBERO , one of the largest open benchmarks for Vision-Language-Action (VLA) policies”
Source E2
Original source quotation: “Meta-World , a premier benchmark for testing multi-task and generalization abilities in robotic manipulation”
Source E3
Original sourceHugging Face — Robotics
NVIDIA has released the GR00T N1.5 model, a cross-embodiment foundation model for generalized humanoid robot reasoning and skills. The model can be fine-tuned using teleoperation data from a SO-101 arm, with a detailed tutorial provided for developers. The release includes instructions for dataset preparation, fine-tuning, evaluation, and deployment.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Environment:
- Not reported
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “This cross-embodiment model processes multimodal inputs, including language and images, to perform manipulation tasks across diverse environments.”
Source E1
Original sourceHugging Face — Robotics
Hugging Face introduces SmolVLA, a 450M parameter open-source Vision-Language-Action model for robotics, trained on community-shared datasets. It outperforms larger models in simulation and real-world tasks, supports asynchronous inference for faster response, and is designed for deployment on consumer hardware.
Evidence and limits
Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.
- Control:
- Not reported
Reported quantities Scroll across to read all columns.| Metric | Value / unit | Basis / context | Evidence |
|---|
| task throughput | 2 other | Basis not reported Source wording: “2× task throughput” | Source E1 |
| task success | 78 percent | Basis not reported Source wording: “78.3% success” | Source E1 |
| task completion time | 9.7 hours | Basis not reported Source wording: “9.7s vs. 13.75s” | Source E1 |
| number of tasks completed | 19 trajectories | Basis not reported Source wording: “19 vs. 9 cubes” | Source E1 |
| number of tasks completed | 78 percent | Basis not reported Source wording: “78% success” | Source E1 |
Dataset details
- License
- Not reported
- Version
- Not reported
- Modalities
- Not reported
- Tasks
- Not reported
- Embodiments
- Not reported
Source excerpts and review record
No manual editorial approval recorded.
Original source quotation: “SmolVLA-450M outperforms much larger VLAs and strong baselines such as ACT on simulation (LIBERO, Meta-World) and real-world tasks ( SO100, SO101 ).”
Source E1
Original source