CoRL 2026|Making Robots More Agile: Starting with Learning to Walk
CoRL 2026|让机器人更会动,从学会「起步」开始
This paper presents LeaP, a learnable source prior for generative robot policies, which replaces fixed standard Gaussian distributions with state-dependent random initialization. LeaP jointly learns the mean and variance of an initial Gaussian distribution, improving action generation in robotic tasks. The method achieves an average success rate of 81.6% in simulation and 80.0% on real robotic arms, outperforming existing baselines like A2A and VITA. The paper is accepted at CoRL 2026, and the code is open-sourced.
Full article
You are reading the complete RoboSignal summary. The publisher’s full article is available at the original source.
Read full article at sourcemp.weixin.qq.com · Opens in a new tab; source language may differ.
This paper introduces LeaP, a learnable source prior for generative robot policies, which improves action generation by incorporating state-dependent random initialization. The method outperforms baselines in both simulation and real-world robotic tasks, demonstrating the effectiveness of modeling uncertainty in action generation.
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
- Environment
- SimulationOpen source S8Open source S13Open source S16Open source S18
Reported numbers
Reported success rate
81.6%
View original evidence
在 RoboTwin 的 15 项仿真操作任务上,其平均成功率达到 81.6 % ,较采用相同编码器与生成器的标准高斯基线提升 25.5 个百分点。
Open source S8Reported success rate
80%
View original evidence
LeaP 平均成功率为 80.0% ,高于 A2A 的 68.3%、VITA 的 56.7% 和 NoPrior 的 46.7%;相较 NoPrior 提升 33.3 个百分点。
Open source S17Reported success rate
85.3%
Reported success rate
91%
View original evidence
在打开笔记本电脑任务的实验中,LeaP 训练 1000 轮达到 91% 成功率,超过四个主要基线训练 3000 轮后的表现,展示了改善收敛效率的潜力。
Open source S15Reported success rate
76.7%
Reported success rate
59.7%
Reported success rate
78%
View original evidence
更关键的是一组三方对照:仅使用预测均值时,成功率为 78.0%;加入固定标准差为 1 的高斯噪声后,反而降至 62.7%;联合学习均值与状态自适应方差则达到 85.3%。
Open source S21success_rate
68.3%
View original evidence
LeaP 平均成功率为 80.0% ,高于 A2A 的 68.3%、VITA 的 56.7% 和 NoPrior 的 46.7%;相较 NoPrior 提升 33.3 个百分点。
Open source S17
- paper: Available Artifact link Open source S10
- dataset: not_reported
What remains unknown
Not established in the collected evidence: Control, Data origin.
Reported performance applies to the described task. It does not establish general autonomy or deployment readiness.
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
在 RoboTwin 的 15 项仿真操作任务上,其平均成功率达到 81.6 % ,较采用相同编码器与生成器的标准高斯基线提升 25.5 个百分点。
Open source S8
实验结果与分析 团队在 RoboTwin 的 15 项双臂操作任务上检验这一设计,涵盖抓取放置、工具使用与双臂协同。
Open source S13
在 Franka Research 3 的抓取方块、关闭盒子、抓取并放置沙袋三项任务中,每项使用 100 条遥操作演示,并随机设置物体位置测试 20 次。
Open source S16
实验在抓取不同瓶子、打开笔记本电脑、交接方块三项 RoboTwin 任务上进行,每项任务评估 100 次。
Open source S18
LeaP 平均成功率为 80.0% ,高于 A2A 的 68.3%、VITA 的 56.7% 和 NoPrior 的 46.7%;相较 NoPrior 提升 33.3 个百分点。
Open source S17
仅用本体感知的 LeaP 平均成功率达到 85.3%,视觉及三种视觉与状态融合方案则为 59.3%—70.3%。
Open source S19
在打开笔记本电脑任务的实验中,LeaP 训练 1000 轮达到 91% 成功率,超过四个主要基线训练 3000 轮后的表现,展示了改善收敛效率的潜力。
Open source S15
同一先验使流匹配的平均成功率从 47.7% 提升至 85.3%;在扩散桥中,相比采用先验均值的确定性起点,完整 LeaP 分布将成功率从 68.7% 提升至 76.7%。
Open source S24
仅通过流匹配损失训练先验,平均成功率为 59.7%,高于无先验基线的 47.7%,但仍低于完整模型的 85.3%。
Open source S22
更关键的是一组三方对照:仅使用预测均值时,成功率为 78.0%;加入固定标准差为 1 的高斯噪声后,反而降至 62.7%;联合学习均值与状态自适应方差则达到 85.3%。
Open source S21
论文标题:Where Should Action Generation Begin? A Learnable Source Prior for Generative Robot Policies 论文链接:https://arxiv.org/abs/2606.17408 项目仓库:https://github.com/SEU-VIPGroup/LeaP 项目主页:https://daimeipo.github.io/LeaP/ 图 1|三种动作生成起点:标准高斯噪声、由观测确定的起点,以及 LeaP 学习的条件概率分布。
Open source S10
Implications for data suppliers
RoboSignal interpretation and collection questions, not statements of buyer demand.
- Confirm the required data type and collection setting with the buyer; this source does not establish a complete collection specification.
- Compare the reported units and scope before using these quantities in a budget. Recording hours, sensor-hours and trajectories are different measures.
- Validate demand and acceptance criteria with a buyer before scaling. Publication, popularity and a research result do not establish a purchase commitment.
Source:Machine Heart — Robotics on WeChat · mp.weixin.qq.com