Skip to content
Source
Chinese robotics — Hardware and sensing·· 4 hours agoSignalEditorial score85

From Data Factory to Training Hub: Can Simulation Platforms Bring a Paradigm Shift to Embodied Intelligence?

从数据工厂到训练中枢,仿真平台能否为具身智能带来范式革命?

Summary

The article explores the evolution of embodied intelligence from data collection to training infrastructure, emphasizing the need for simulation platforms that enable continuous learning and iteration. It discusses the limitations of raw data and the importance of creating a closed-loop system for model training and real-world deployment, with a focus on physical simulation and the integration of real-world feedback.

Source: Chinese robotics — Hardware and sensing · Read original article ↗

Article text · Machine translation into English

谋先飞孙家琦:从数据工厂到训练中枢,仿真平台能否为具身智能带来范式革命?

In the past two years, the focus of the embodied intelligence industry has been continuously moving forward.

Initially, the market concentrated on whether robots could perform more complex actions; as VLA, world models, and end-to-end control rapidly developed, the industry began to attribute the bottleneck to "not enough data." True machine data collection, internet videos, first-person devices, motion capture, and synthetic data became different directions that various teams sought for incremental gains.

But when robots truly entered the training and industrial implementation phase, the issue was no longer just "where to get the data," but whether data from different sources could be transformed into physical experiences that robots could execute, verify, and sustainably reuse; whether failures encountered during true machine deployment could be fed back into the training system to form the next round of iterative tasks.

This means that embodied intelligence is moving from a simple data competition into a competition for training infrastructure.

During the 2026 World Robotics Conference, Motphys Co-founder and Chief Architect Sun Jiaqi, at the "2026 World Robotics Conference · Embodied Intelligence Commercialization Forum" hosted by Y欧 proposed that the positioning of simulation needs to be upgraded from a "data factory" to a "training hub."

谋先飞孙家琦:从数据工厂到训练中枢,仿真平台能否为具身智能带来范式革命?

Simulation is not just about generating more data for models, but providing an interactive, intervenable, verifiable, and continuously learnable computational world, connecting real data, simulation data, strategy training, capability evaluation, and real machine deployment.

This is also the embodied intelligence simulation training infrastructure that Motphys is currently building.

Data is abundant, but reusable experience for robots remains scarce.

Embodied intelligence does not lack data sources.

True machines can generate data that is closest to real-world deployment, internet videos contain a large amount of human behavior experience, Ego devices and motion capture systems can record first-person information and human motion trajectories, and simulation can low-costly expand scene and task scales.

However, using these whether true machines or real sensors to collect data has one problem: real sensors during data collection are affected by time synchronization errors, device aging, heating and environmental changes, etc. The data formats, coordinate systems, sensor configurations and sampling frequencies of different devices are also different, making it difficult to directly form a unified, structured, and reusable dataset.

More importantly, human behavior data cannot be directly executed by robots. A video of a human opening a drawer can record hand positions and motion trajectories, but cannot directly tell a robot the damping of the drawer, how the contact relationship changes, how much torque is needed, and how to adjust actions when sliding or jamming occurs. Therefore, converting human experience into robotic capabilities cannot only be imitation at the visual, pose, or trajectory level, but also requires understanding of task intent, physical state transitions, contact topology, and torque constraints.

Traditional methods can complete part of spatial position and pose alignment, but struggle to simultaneously ensure consistency in task results and physical interactions. Retargeting requires an interactive, computable physical environment to align, clean, and validate data from true machines, videos, motion capture, and different sensors.

From this perspective, embodied intelligence truly lacks not just data, but physical experience that can be reused by different robot bodies.

The value of simulation lies not only in generating data, but in forming a training loop

The traditional training process for embodied models usually involves data collection, cleaning and labeling, offline training, and real machine testing.

In this linear process, training and evaluation are often disconnected. After a model fails on a real machine, the team needs to re-locate the problem to determine whether it comes from data, model, physical parameters, or control strategy, and then manually organize the next round of training.

The change brought by simulation is to turn this process into a continuous iteration loop.

After each round of model training produces a new Checkpoint, it can quickly complete task set evaluation in a simulation environment to determine whether the model's capabilities have improved, whether there is forgetting, and in which scenarios and physical conditions the model is prone to failure. These failure cases can be further converted into new scenarios, parameters, and training tasks for the next round of model training.

The significance of simulation evaluation is not only to externally refresh a Benchmark score. For model development teams, the longer-term value is to establish an internal unified, repeatable, and continuously accumulating evaluation system.

As models continue to iterate, simulation assets, task definitions, environment parameters, failure cases, and evaluation standards will become reusable infrastructure. Errors generated from real machine deployment can also be fed back into the simulation environment, continuously calibrating physical parameters and strategy robustness.

When simulation is connected to data generation, strategy training, model evaluation, real machine deployment, and error feedback, it is no longer a single-point tool, but the training hub for robots.

The true training world must be able to compute physical results

Currently, world models and spatial intelligence have become important directions in the field of embodied intelligence. However, in Moxianfei's view, generating a visually realistic three-dimensional world does not equal to building a world that can train robots.

Robots need to perform actions in the environment and obtain the physical results of those actions.

When a mechanical arm grasps an object, a dexterous hand operates a tool, or a humanoid robot transports a box, the model needs to handle not only spatial positions, but also contact, friction, collision, joint constraints, flexible deformation, torque feedback, and changes in object states.

Therefore, a truly valuable virtual world for training must meet several conditions: it must be interactive, computable, able to actively change key variables, able to validate task results, and able to transfer strategies to real robots.

This is the core difference between physical simulation and pure visual generation and three-dimensional content production.

For Contact-rich tasks, even minor physical deviations can be amplified on real hardware. Tasks such as robotic arm grasping, dexterous hand manipulation, precision assembly, and cable insertion involve continuous, multi-point, and subtle contact processes, which impose higher requirements on the accuracy and stability of the physics solver.

Moxianfei has long focused on the underlying physics solver, covering rigid body, flexible body, cloth, rigid-flexible coupling, contact, friction, joint, and tactile sensor simulation capabilities, aiming to enable robots to gain physical experience that can be transferred to the real world.

From physics engine to data and training, Moxianfei builds end-to-end capabilities

Moxianfei was founded in 2021, positioning itself as a company providing simulation training infrastructure for embodied intelligence.

The company builds MotrixSim simulation platform, MotrixLab imitation and reinforcement learning training framework, and Mot,rixGen embodied intelligence data engine based on its self-developed physics solver, forming a product system covering simulation environment, data generation, strategy training, capability evaluation, and real-world migration.

MotrixSim mainly addresses issues such as physical consistency, complex contact, high-concurrency simulation, and Sim2Real migration in robot training, and can be integrated with different bodies such as humanoid robots, quadruped robots, robotic arms, and dexterous hands.

MotrixLab integrates common imitation learning and reinforcement learning frameworks, supporting CPU/GPU heterogeneous computing, enabling robots to perform large-scale parallel trial-and-error and strategy training in simulation environments.

MotrixGen focuses on data generation and data augmentation for embodied intelligence, expanding long-tail, suboptimal, and failure samples that are difficult to collect at low cost in the real world through task definition, expert trajectory synthesis, and domain randomization.

These three layers of capabilities are not independent product modules, but collectively form a Real→Sim→Real loop.

After real-world scenes, assets, and physical parameters are imported into the simulation environment, the system can generate and enhance training data, complete strategy training through imitation learning or reinforcement learning, and then migrate the strategy to real robots. Errors that occur during real-world execution can be fed back into the simulation environment to calibrate asset parameters and training distributions.

Moxianfei aims to establish a data flywheel that enables continuous training, evaluation, and iteration.

The key to Real2Sim is making assets 'interactable'

Real2Sim is often understood as scanning or reconstructing real-world scenes into a virtual environment. However, for robot training, visual and geometric reconstruction is just the first step.

A SimReady asset that can be used for robot training also needs to have computable physical properties.

For example, a drawer asset not only needs to include appearance, size, and sliding joints, but also needs to describe the force-displacement curve during the pulling process; a refrigerator door needs to align with the opening angle and torque change; components with magnetic attraction or sticking structures need to replicate the resistance changes in different stages.

These parameters need to be entered into the simulation environment through real measurement, data fitting, and physical parameter identification, so that the robot's interaction in the virtual environment remains consistent with the real physical process.

Moxianfei currently supports two types of Real2Sim paths.

For clear industrial deployment scenarios, high-precision replication of scenes and assets can be achieved through video, multi-angle images, scanning, and real physical measurements.

For pre-training and benchmark construction, parameterized generation can be used, combining different assets, layouts, physical parameters, and interaction mechanisms through digital labels to build a larger-scale task distribution.

The former emphasizes accurate recreation of real scenes, while the latter emphasizes scalable generation and compositional generalization. Both serve different stages of robots from pre-training to real deployment.

Training data cannot only consist of successful samples

Another important value of simulation is the ability to systematically generate data that is difficult to collect in the real world.

In real environments, large-scale collection of failure data typically means higher equipment wear, safety risks, and labor costs. However, for robot models, learning only successful trajectories is insufficient to form stable generalization capabilities.

Therefore, Moxianfei does not only collect optimal expert trajectories in simulation data generation, but also generates suboptimal trajectories, failure trajectories, and trajectories close to the success or failure boundary through perturbation and Planning algorithms.

Simulation can simultaneously launch hundreds, thousands, or even more parallel environments, enabling systematic exploration of different parameters, states, and failure conditions through computing power, without being limited by real physical timing and hardware quantity.

In the parts of quadruped motion, full-body control, mechanical arm grasping, and dexterous hand operation tasks that MoXianFei has demonstrated, strategies can complete simulation training without the need for real-world demonstration data and migrate to real robots.

This does not mean that simulation can completely replace real-world data. Real data is responsible for providing calibration benchmarks, while simulation is responsible for expanding training distribution and covering long-tail conditions. Real-world validation brings new errors and issues back to simulation.

Real data and simulation data are not in a substitution relationship, but rather an amplifying relationship.

Robot training cannot be bound to a single hardware ecosystem

Embodied intelligence training is a complex systems engineering that includes physical simulation, data generation, strategy training, parameter synchronization, model evaluation, and real-world deployment.

Different stages have different requirements for CPU, GPU, cloud computing power, and edge devices. If the underlying training infrastructure is deeply bound to a single hardware system, it may limit its application in industrial, localized deployment, and security-sensitive scenarios.

Therefore, heterogeneous training architecture is an important part of MoXianFei's product roadmap.

While being compatible with the NVIDIA ecosystem, MoXianFei is also advancing operator-level adaptation for AMD, Apple, and domestic CPUs and GPUs. It has already completed dual certification for Huawei Kunpeng Compatible and Native, and is currently advancing adaptation for domestic computing power platforms such as TianShuZhiXin and Moor Threads.

The rapid development of NVIDIA Isaac, Newton, and other ecosystems is proving that robot simulation training has become a global infrastructure track. MoXianFei's chosen path is not simply a single-point performance comparison with international platforms, but building a self-controlled, heterogeneous-adaptive, and deployable simulation training system in the context of China's robotics industry.

This system also needs to connect the robot body, sensors, actuators, computing hardware, algorithm frameworks, and developers, forming an open embodied intelligence simulation ecosystem.

From simulation tools to embodied intelligence training infrastructure

Embodied intelligence is moving from concentrated demonstrations of its capabilities and technical demos to model training, engineering validation, and industrial deployment phases.

The core issue in this phase is no longer whether the robot can perform a certain action, but whether the model can stably reproduce tasks, whether it can migrate to different bodies and environments, and whether deployment failures can be quickly converted into new training experiences.

This is also the reason why MoXianFei is moving from a simulation platform to a training infrastructure.

Currently, MoXianFei has jointly released an embodied intelligence simulation training closed-loop solution with Digua Robotics and Yingmou Technology, connecting 3D scene generation, SimReady assets, physics simulation, strategy training, and computing toolchains into the same workflow; at the same time, it is collaborating with the National and Local Jointly Built Humanoid Robot Innovation Center to explore a true machine and simulation integrated data system of 'true machine data collection—simulation recreation—data augmentation—closed-loop iteration'.

In MoXianFei's view, the next stage of competition in embodied intelligence will not only occur at the level of robot bodies and models, but also at the level of data production, physical simulation, training evaluation, and real machine migration infrastructure.

True machines provide real feedback, simulation completes large-scale trial and error and data expansion, and the training framework converts deployment failures into the next round of capabilities. Forming a Real-to-Sim closed loop, this also includes the realization of video rapid generation of interactive simulation assets; data augmentation, which uses a small amount of high-quality data as a basis to achieve massive automated expansion; and even actively adding disturbances to adapt to real hardware defects, solving the problem of the gap between virtual and real migration. The more complete this closed loop is, the more likely robots are to move from one-time demos to continuous learning and large-scale deployment.

From data factory to training hub, the role of simulation is changing.

What MoXianFei hopes to build is a set of embodied intelligence training infrastructure that allows robots to 'train in the virtual world, take effect in the real world, and continuously evolve through real feedback.'

Source:Chinese robotics — Hardware and sensing · timeline.sohu.com

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).