How to evaluate robot demos and benchmark claims
RoboSignal · Published · Version 1.0
A robot demo shows an observed behavior under a particular setup. A benchmark adds a stated evaluation procedure. Neither label alone establishes unattended reliability or deployment readiness. This checklist helps readers keep the reported result, its test conditions and its unanswered questions together.
1. Identify the tested system
Record the robot hardware, controller or model version, camera arrangement and task. Determine whether the shown system matches the released checkpoint or product. Ask which parts were changed for the test. OpenVLA provides a concrete project where model adaptation and evaluation need to be read together. Do not compare results from differently configured systems as if only the model name changed.
2. Separate training, tuning and testing
A test claim should specify what was held out. A new object, location, robot platform and instruction are different kinds of generalization. Record any target-task fine-tuning, calibration or task-specific prompt selection. Inspect the baseline’s conditions too. Fairness depends on the actual protocol, not on whether a comparison chart uses matching colors. If the split is not reported, leave it unknown.
3. Keep the denominator and failures
A success rate needs successful attempts and the total number of attempts, together with a success rule. Ask whether failed starts, human rescues, retries and timeouts count. A highlight reel cannot provide a trial denominator. When only a rate is reported, preserve it as a source-reported result and state that the count is missing. Avoid precision that the source cannot support, and do not pool tests with incompatible definitions.
4. Describe control and assistance
Record continuous teleoperation, shared control, high-level human choices, interventions and resets separately where possible. Use not reported when the source is silent. Training with teleoperation and executing a learned policy are different stages. Likewise, distinguish simulation from physical trials. A useful demo can still require substantial assistance; document that condition rather than turning it into an unsupported autonomy label.
5. Trace the claim to its evidence
Use the paper, project report or test protocol for the exact claim. A vendor’s blog, paper and social post can represent one origin even when they are three URLs. Independent confirmation requires a genuinely separate source of evidence. Identify who ran the evaluation and who selected the examples. Repeated reporting can establish attention without establishing scientific merit or truth.
6. Turn the result into a bounded conclusion
State what was tested and what the evidence supports under that setup. Leave broader safety, endurance, economics and deployment claims open unless separately established. Download the evaluation checklist CSV to capture each condition and its source. Follow robotics research for source-linked developments. RoboSignal’s automatic publication checks do not substitute for human factual calibration, which remains in progress.
Primary references
- OpenVLA: project, paper and evaluation
- ACT: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- MuJoCo: official simulation overview
Reference links checked 2026-10-04. Project claims remain attributed to their original source. This page is not a certification or a live test of the referenced system.
Related reading
Cite this reference
RoboSignal. “How to evaluate robot demos and benchmark claims” (2026-10-04), version 1.0. https://robosignal.ai/resources/robot-evaluation-checklist. Cite the original project separately for its own reported results.