DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
Overview
DexPolicy introduces a reinforcement learning method that manages exploration scale during training to improve dexterous manipulation. It is tested across multiple policy optimization scenarios.
Generated from attributed reports · 2 days agoUpdated
Event evidence and corrections
0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.
Reported quantity · mean deterministic Target success: 49.4 percent · Basis not reported · Differing source assertions
“mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO).”
Exact source · revision 1Source owner not reported
Reported quantity · mean deterministic Target success: 14.1 percent · Basis not reported · Differing source assertions
“mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO).”
Exact source · revision 1Source owner not reported
Reported quantity · mean deterministic Target success: 32 percent · Basis not reported · Differing source assertions
“mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO).”
Exact source · revision 1Source owner not reported
Reported quantity · mean deterministic Target success: 25 percent · Basis not reported · Differing source assertions
“mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO).”
Exact source · revision 1Source owner not reported
Reported quantity · mean deterministic Target success: 10 percent · Basis not reported · Differing source assertions
“mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO).”
Exact source · revision 1Source owner not reported
Reported quantity · mean deterministic Target success: 8.3 percent · Basis not reported · Differing source assertions
“mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO).”
Exact source · revision 1Source owner not reported
Report timeline
Follow attributed reports and material updates.
- arXiv Robotics — research abstractsDexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
This research presents DexPolicy, a reinforcement learning approach that explicitly manages exploration scale during training to enhance dexterous manipulation. The method is tested across multiple policy optimization settings (PPO, GRPO, FPO) and shows significant improvements in task success rates, with FPO achieving the highest performance on both simulated and real-world robotic arms.
Event attention history
Current attention 3·Peak within the comparable range 9(2026-10-02 06:00 UTC)·Change within the comparable range over 24 hours -49%
The trend compares participants observed continuously; its scope may be narrower than current attention totals. Hover or click to explore hourly observations, or use the left and right arrow keys.