ICRA 2026 | First 36 Degrees of Freedom Dual-Arm Dexterous Manipulation VLA Model, Open Source!
ICRA 2026 | 首个36 自由度双臂灵巧操作VLA 模型,开源!
Dexora is the first open-source VLA model designed for dual-arm, high-degree-of-freedom (36 DoF) dexterous manipulation, addressing the limitations of previous models that could not simultaneously handle bimanual coordination and fine-fingered tasks. It combines real-world and simulated data with a quality-aware training framework to improve robustness and generalization.
Source: Robotics — Chinese web discovery · Read original article ↗
Article text · Machine translation into English
On this page
Sina Tech Stock

(Source: LeFeng Network)
Original author: Official account "Shenlan Embodied Intelligence"
Original link: https://mp.weixin.qq.com/s/yGE3tLQywqf4wvsOlj6WxA
Supports dual-arm coordination, dual dexterous hands with high degrees of freedom for fine operations
—— End-to-end VLA model
Previously, mainstream VLA systems either focused on dual-arm low-degree-of-freedom gripper control or specialized in single-arm high-degree-of-freedom dexterous hand operations, always failing to meet the dual demands of dual-arm coordination and fine finger actions.
For example, inserting a piston requires precise dual-arm coordination, while tasks like screwing a bottle cap or fine object retrieval depend on multi-finger flexible control, and such tasks have long lacked a unified VLA solution.
Recently, Dexora, the first open-source VLA model natively designed for dual-arm dual high-degree-of-freedom (36 degrees of freedom) dexterous operations, has broken through the previous limitations of VLA, providing a new paradigm for the practical application of general dexterous robots.
Hardware and Teleoperation: A Data Foundation for Virtual-Reality Synergy
The core prerequisite for high-degree-of-freedom dexterous operations is precise and scalable data collection. Dexora discards a single teleoperation solution and creates a "exoskeleton + VR" hybrid teleoperation system, simultaneously driving physical robots and MuJoCo digital twins, addressing the root issues of data collection precision and scale.
▲Dex hardware and hybrid teleoperation system diagram ©【Shenlan Embodied Intelligence】translation
Can achieve twisting, pinching, and other human-like fine operations, with a total of 36 degrees of freedom, providing the hardware foundation for complex dual-hand coordination tasks.
Custom exoskeleton backpack captures the overall movement of the robotic arm (no drift, low latency), Apple Vision Pro enables markerless finger tracking, balancing the stability of large-range arm movement with the flexibility of fine finger actions.
▲Simulation / Real Data Set Object and Task Distribution Diagram ©【Shenlan Embodied Intelligence】translation
More importantly, the design of virtual-real synchronization:
Teleoperation commands are simultaneously sent to the physical robot and the digital twin, with both sensor data (4 RGB, 36 degrees of freedom joint states) recorded at 20Hz synchronously.
This ensures the authenticity of real-world data while enabling low-cost expansion of task scenarios in simulation, forming a "real + simulation" complementary dataset.
The final constructed dataset includes 100,000 simulation trajectories (6.5 million frames), 10,000 real teleoperation trajectories (2.92 million frames), covering 30 types of simulated objects and 17 types of real objects, balancing basic grasping and fine dexterous tasks.
Model Architecture: Diffusion Transformer + Quality Discriminator Dual Core
Dexora is based on a data quality-aware end-to-end VLA architecture, using a diffusion Transformer to generate actions and an offline quality discriminator to filter noisy data, solving the industry pain points of inconsistent teleoperation data quality and unstable high-dimensional action training.
▲Dexora Overall Overview Diagram (Motivation / Data / Architecture / Performance / Generalization) ©【Shenlan Embodied Intelligence】translation
Diffusion Transformer Strategy Network
A decoder-only diffusion Transformer is used as the core strategy, taking multi-view RGB images, language instructions, and current joint states as input, and outputting a 36-degree-of-freedom continuous action sequence.
The model uses T5 to encode language instructions, SigLip to encode image features, and alternately injects Transformer blocks, generating actions through denoising prediction, balancing multi-modal fusion capabilities and high-dimensional action modeling accuracy.
Data Quality Discriminator
Teleoperation data often contains low-quality trajectories due to operational errors and sensor noise, and direct training can lower model performance.
Dexora designs an offline discriminator that uses "motion smoothness + task success rate" as dual standards to filter data:
The discriminator, using a frozen pre-trained strategy as a benchmark, predicts trajectory quality scores (0-1), and during training, converts the scores into weights, with high-quality trajectories having high weights and low-quality trajectories having low weights, with the formula simplified as:
Where is the quality weight, is the predicted noise, and is the real noise. This design allows the model to focus on effective data, significantly improving the training stability of high-dimensional dexterous actions.
Training Process: Three-Stage Progressive Optimization
Dexora adopts a "simulation pre-training + discriminator training + real fine-tuning" three-stage training, balancing basic capabilities and dexterous skills, achieving smooth migration from simulation to reality.
▲Dex Data Filtering, Discriminator Training, Quality-Aware Training Framework Diagram ©【Shenlan Embodied Intelligence】translation
First stage: 100,000 simulation trajectories pre-training, allowing the model to master basic operations such as grasping and assembly, forming initial action generation capabilities;
Second stage: using the filtered high-quality real data to train the discriminator, enabling it to accurately identify trajectory quality;
Third stage: using all real data for fine-tuning the strategy, guiding the model from basic capabilities to advanced dexterous skills such as screwing bottle caps and fine object retrieval through quality weights.
Performance and Generalization
Experimental results show that Dexora achieves breakthroughs in three dimensions: basic tasks, dexterous tasks, and cross-form generalization, validating the effectiveness of the dual-arm dual high-degree-of-freedom design and quality-aware training.
▲Basic Task Example Diagram ©【Shenlan Embodied Intelligence】translation
▲Basic Task Success Rate Comparison Table ©【Shenlan Embodied Intelligence】translation
Average success rate of 89.6%, with 7 out of 12 tasks having success rates exceeding 90%, showing significant advantages in dual-arm coordination tasks (such as handing objects with both hands, separating nested bowls), far surpassing GR00T N1 (82.1%) and π0 (50.4%) baselines.
▲Dexterous Task Example Diagram ©【Shenlan Embodied Intelligence】translation
▲Dexterous Task Success Rate Comparison Table ©【Shenlan Embodied Intelligence】translation
Average success rate of 66.7%, an increase of 15 percentage points compared to the best baseline GR00T N1 (51.7%), especially in tasks requiring dual-hand coordination and multi-finger control such as screwing a bottle cap and fine dough manipulation, where baselines almost fail, but Dexora can still complete them stably.
▲Out-of-Distribution Generalization Performance Diagram ©【Shenlan Embodied Intelligence】translation
First, out-of-distribution generalization, with success rates only slightly decreasing in unknown backgrounds, lighting, objects, and occlusions, showing strong robustness;
Second, cross-form migration, the 36-degree-of-freedom model can directly adapt to single-arm grippers, dual-arm grippers, and single-arm low-degree-of-freedom hands, without the need for retraining, only requiring adaptation of action dimensions, breaking the form dependency of VLA.
▲Basic Task Success Rate Comparison Table ©【Shenlan Embodied Intelligence】translation
Ablation experiments show that the quality discriminator can reduce action jitter and improve task stability, proving that "real data + quality awareness" is a core element of dexterous VLA.
▲Quality Discriminator Effect Comparison Trajectory Diagram ©【Shenlan Embodied Intelligence】translation
Dexterous VLA: Value and Limitations Coexist
Dexora is the first native dual-arm dual high-degree-of-freedom VLA, proving that high-degree-of-freedom models can be downward compatible with low-degree-of-freedom devices, providing a new approach for general robots to "high-dimensional training, low-dimensional deployment."
Virtual-real synergy data collection + quality-aware training solves the problems of scarce and noisy dexterous data, providing a reference for high-dimensional VLA data construction.
▲Dexora and Mainstream VLA Morphological Coverage Comparison Diagram ©【Shenlan Embodied Intelligence】translation
At the same time, there are existing limitations:
Hardware dependency: The 36-degree-of-freedom system has high costs and is difficult to popularize quickly, and there is no tactile feedback, resulting in low success rates for contact-sensitive tasks such as screwing a bottle cap.
Task limitations: Inability to handle complex long-time sequence tasks (such as multi-step assembly) and dynamic environment adaptation.
Generalization boundaries: Stability in cross-material and extreme scenarios needs improvement.
Previously, VLA systems either "can use both hands but not finely" or "can be fine but not use both hands."
Dexora is the first to unify both, using an open-source model to lower the threshold, providing feasible solutions for service robots and industrial dexterous operations scenarios.
In the future, if tactile feedback and long-time sequence reasoning can be integrated, Dexora has the potential to further narrow the gap with human operations, promoting general dexterous robots from the lab to real-world applications.
Ref
Paper Title: Dexora: Open-source VLA for High-DoF Bimanual Dexterity
Paper Authors: Zongzheng Zhang, Jingrui Pang, Zhuo Yang, Kun Li, Minwen Liao, Saining Zhang, Guoxuan Chi, Jinbang Guo, Huan-ang Gao, Modi Shi, Dongyun Ge, Yao Mu, Jiayuan Gu, Rui Chen, Hao Dong, Huazhe Xu, Li Yi, Yixin Zhu, Hang Zhao, Pengwei Wang, Shanghang Zhang, Guocai Yao, Jianyu Chen, Hongyang Li, Hao Zhao
Paper Link: https://arxiv.org/pdf/2605.18722
Project Link: https://dexoravla.github.io/
LeFeng Network
Loading...
Video Live Beautiful Photos Blog Highlights Government Humor Gossip Emotion Travel Buddhism Crowdsourcing
Source:Robotics — Chinese web discovery · finance.sina.cn