REDIRECT: A 1% Fix for Bad Robot Habits
REDIRECT: A 1% Fix for Bad Robot Habits
Robots can develop bad habits from a few defective moments in otherwise useful teleoperation. This paper introduces REDIRECT, a method that fixes these local differences using only 1% of the full-training sample budget. By localizing problematic behavior and assigning coherent continuations, REDIRECT improves success rates from 68.2% to 90.3% across three ManiSkill tasks and recovers most of the gap compared to clean-retraining.
Source: arXiv Robotics — research abstracts · Read original article ↗
Article text · Original source · English
arXiv:2610.03997v1 Announce Type: new Abstract: Robots can acquire bad habits from a few defective moments in otherwise useful teleoperation. In mixed-quality robot data, normal and problematic demonstrations share most task behavior and differ only at a local action continuation. Full retraining is costly, while fine-tuning on clean data alone offers limited recovery under a small update budget. We ask whether the local difference can instead be fixed with only 1% of the full-training sample-backward budget. Using only episode-level retained/problematic labels, REDIRECT localizes the branch, assigns a coherent retained continuation to the original observations around it, and anchors shared behavior, without frame-level annotations or additional interaction. Across three randomized ManiSkill tasks and three seeds, REDIRECT raises the mean success rate from 68.2% to 90.3%, recovering 88.8% of the clean-retraining gap versus 22.1% for matched-compute fine-tuning. On a PiPER arm, it recovers 86.7-92.9% of the clean-only gap across cup insertion and towel folding. Local robot habits can therefore be repaired by spending the update budget on the behavioral difference rather than relearning shared behavior.
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
Reported numbers
Reported success rate
68.2%
Reported success rate
90.3%
What remains unknown
Not established in the collected evidence: Environment, Control, Data origin.
Reported performance applies to the described task. It does not establish general autonomy or deployment readiness.
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
s the mean success rate from 68.2% to 90.3%, recovering 88.8% of the clean-retraining gap versus 22.1% for matched-compute fine-tuning. On a PiPER arm, it recovers 86.7-92.9% of the clean-only gap across cup insertion and towel folding. Local robot habits can therefore be repaired by spending the update budget on the behavioral difference rather than relearning shared behavior.
Open source S4
Source:arXiv Robotics — research abstracts · arxiv.org