Robotics dataset evidence checklist
RoboSignal · Published · Version 1.0
Evaluate a robotics dataset by what its documentation establishes for your task. Start with the original release, preserve its units, and record missing fields as not reported. This checklist helps compare evidence without treating scale, download access or a project name as a quality certificate.
1. Identify the release and its origin
Record the dataset name, release version, original project URL and the date you checked it. Describe whether it contains physical robot interactions, human egocentric recordings, simulation or a mixture. Preserve those types when comparing data. DROID is a robot manipulation dataset reference; Ego4D is an egocentric video reference. These are different evidence categories, not substitutes for one another.
2. Preserve the scale basis
Copy each reported value together with its unit, scope and source location. Check whether hours represent unique elapsed collection time, summed recordings or another basis. Camera count alone cannot tell you the extent of overlap. Do not convert trajectories to hours without a duration source. If numbers refer to different release versions, keep separate rows. The local calculator can explain a fully overlapping multi-camera setup; it cannot infer the timing of an unknown dataset.
3. Inspect learning compatibility
Record robot platforms, observation modalities, action representation, timestamps, task labels and delivery format. Identify the intended training method and the loader it expects. LeRobot documentation is one practical entry point for inspecting robot-learning workflows. A project supporting a format does not mean every release is ready for that workflow. Check a small permitted sample and record gaps before assuming compatibility.
4. Separate access from rights
A working download link establishes access to that object at that time. It does not establish commercial permission, redistribution permission or rights to every underlying component. Follow the exact release terms and any application process. Record licence text, version and the relevant URL. Leave ambiguous conditions unresolved and ask the publisher where needed. This checklist records documentation; it is not a legal interpretation or consent certification.
5. Examine coverage and splits
Ask which tasks, objects, sites, people and robot embodiments the data covers. Record how train, validation and test partitions were made. For a new-environment claim, inspect environment separation; for a new-object claim, inspect object separation. Closely related sequences should not silently become independent tests. Note whether failures are included and what was excluded during cleaning. Missing evidence is an actionable follow-up question, not an automatic negative score.
6. Keep an evidence trail
Download the evidence checklist CSV. Each field has a question and an empty slot for the release’s answer, exact source location and check date. Keep quoted facts distinct from your interpretation. Save the release identity so later corrections can be traced. Use the dataset reference map to start at primary projects, and Data & Training coverage to follow new announcements.
Primary references
- DROID: official dataset project
- Ego4D: official project and access information
- LeRobot: official documentation
Reference links checked 2026-10-04. Project claims remain attributed to their original source. This page is not a certification or a live test of the referenced system.
Related reading
Cite this reference
RoboSignal. “Robotics dataset evidence checklist” (2026-10-04), version 1.0. https://robosignal.ai/resources/robotics-dataset-evidence. Cite the original project separately for its own reported results.