Skip to content
Source
Chinese robotics — Dataset and collection discovery·· 1 hours agoSignalEditorial score85

Embodied-AI-Guide: A Comprehensive Tutorial for Embodied AI

具身智能技术指南Embodied-AI-Guide

Summary

The Embodied-AI-Guide is a detailed tutorial and resource hub for understanding and practicing embodied AI. It covers foundational concepts, practical implementation, and advanced topics such as reinforcement learning, world models, and simulation-to-reality transfer. The guide includes a step-by-step tutorial for training and evaluating an operation policy using RoboTwin 2.0, along with curated resources for further learning and research.

Editorial context

The Embodied-AI-Guide is a comprehensive resource for newcomers to embodied AI, offering a structured learning path that combines practical projects with theoretical overviews. It emphasizes the challenges of data collection, strategy design, and performance evaluation in embodied systems, while also providing tools and frameworks for implementation.

Source: Chinese robotics — Dataset and collection discovery · Read original article ↗

Article text · Machine translation into English

On this page

国内最热门的具身智能技术指南
A compendium-style Chinese knowledge base and resource index for embodied intelligence

From Here · Hands-on Learning · Cognitive Resources · Algorithm Chapter · Infrastructure Chapter · Control Chapter · Hardware Chapter

GitHub repo stars  Visitors  PRs Welcome  License

📚 This project aims to help newcomers quickly build domain awareness: using a practical project to guide hands-on entry into embodied intelligence, while organizing the main technologies involved in embodied intelligence in an encyclopedia format, so that readers can clearly understand what problems different technologies can solve and have a clear direction for further exploration.

Welcome Star / Share / Submit PR. For collaboration and communication, please email [email protected], or add project initiator WeChat TianxingChen_2002 (please note: institution + name + purpose).

📢 News|Project Progress

📝 2026-01-15: Embodied-AI-Guide completed content reorganization

⭐️ 2025-12-18: GitHub Stars exceeded 10,000

❤️ 2025-03-15: Embodied-AI-Guide officially open-sourced

🧑‍💻 Related Projects|Related Open-Source Projects

⭐️ Lumina Call (Embodied Intelligence Recruitment): website

⭐️ Datawhale every-embodied (Embodied Intelligence Tutorial): repo

🦉 Lumina Embodied Intelligence Community

Community Homepage: lumina-embodied.ai. Scan the QR code below to join the discussion group; if the QR code is invalid or you cannot join the group, please email [email protected], or add the project initiator's WeChat TianxingChen_2002 (please note: institution + name + purpose).

Lumina 具身智能社区

🐣 (1) Start From Here —— From Here

Embodied intelligence refers to an intelligent system that operates based on physical entities, acquiring information, understanding problems, making decisions, and taking actions through interaction with the environment, thereby generating intelligent behavior and adaptability.

Put more plainly: large models are already very good at handling text and images, because there is a massive amount of ready-made data on the internet; but to get a robot to actually reach out and pick up a cup on the table, the model must output continuous, physically consequential actions, and such data cannot be scraped from the internet — each piece must be collected by someone using a real device, and mistakes can break things. This 'where does the data come from, how to represent the actions, and how to measure performance' trilogy of questions forms the core contradiction of embodied intelligence today, and is the main thread that repeatedly appears in this guide.

(1.1) How —— How to Use This Guide

The design philosophy of this project is 'one main thread + one panoramic view':

  • One main thread: Use Chapter 2 hands-on tutorial to run through a complete workflow of an operational strategy within a week;
  • One panoramic view: Use Chapters 3 ~ 7 in an encyclopedia format to cover algorithms, infrastructure, control, and hardware, helping you determine what problems each technology solves and whether it's worth delving into.

Recommended reading order: First run through Chapter 2, then search for subsequent chapters based on your interests, there's no need to read them in order.

(1.2) Learning Path Recommendations (Choose Starting Point Based on Background)

Readers with different backgrounds don't need to start from the same place, and can refer to the table below to choose a starting point:

Your Background Recommended Starting Point Recommended Path
Complete beginner / Undergraduate student Chapter 2 Hands-on Tutorial (2) Run through the workflow → (3) Build cognitive understanding → (4) Algorithm Chapter → Dive into based on interest
Have CV / NLP / Deep Learning Background Chapter 4 Algorithm Chapter (4) Focus on Robot Learning and VLA → (5) Infrastructure → (2) Hands-on Verification
Have Automation / Robotics Background Chapter 4 Algorithm Chapter (3) Understand the layout → (4) Complement learning and decision-making → (5) Get familiar with simulation and data ecosystem
Bias towards Hardware / Embedded Systems Chapter 7 Hardware Chapter (7) Hardware Chapter → (6) Control Chapter → (2) Hands-on Tutorial
Want to quickly understand the industry and topics Chapter 3 Cognitive Resources (3.1) Methodology → (3.5) Paper List → (3.6) Annual Trends

(1.3) Term Quick Reference Table

Frequently encountered concepts when reading papers and this guide, establish a basic understanding first, details are left for corresponding chapters.

Expand Term Quick Reference Table (20 terms)

Term One-sentence Explanation
Manipulation / Locomotion Operation (changing the environment with arms) and movement (moving the body with legs or wheels), two main threads of embodied intelligence
Policy (Strategy) Mapping from observation to action, which is what is commonly referred to as the 'model'
IL / BC (Imitation Learning / Behavior Cloning) Learning actions from human demonstration data in a supervised manner
RL (Reinforcement Learning) Optimizing the strategy through interaction with the environment and reward signals
VLA (Vision-Language-Action) End-to-end model that maps images and language instructions directly to robot actions
VA (Vision-Action) Strategy that uses only vision and does not accept language instructions; more lightweight and faster when the task is fixed
World Model (World Model) Model that learns 'current state + action → next observation', can be used for mental simulation
WAM (World-Action Model) A class of strategies that first use the world model to predict what will be seen next, then extract actions from the prediction
Teleoperation (Teleoperation) Human controls the robot through joysticks, motion capture, or master-slave devices, which is the main means of data collection for real devices
Demonstration (Demonstration / Trajectory) A complete record of a task execution, usually includes observations, actions, and timestamps
Sim2Real Gap Differences between simulation and the real world in terms of physics, rendering, and noise, which can lead to performance drop when transferring strategies
Real2Sim Reconstructing real scenes and objects into simulation to reduce the Sim2Real Gap
Affordance (Affordance) The areas or ways an object can be operated on, such as a handle can be grabbed, a button can be pressed
DoF (Degrees of Freedom) The number of joints that a robot can move independently
End-effector (End-effector) The execution component at the end of a robotic arm, such as a gripper, suction cup, or dexterous hand
FK / IK (Forward Kinematics / Inverse Kinematics) Calculating the end-effector pose from joint angles / Solving joint angles from the end-effector pose
URDF XML format describing robot links, joints, and inertia, which is a common input for simulation and control
MPC (Model Predictive Control) Generating control inputs by repeatedly solving optimization problems within a rolling horizon
WBC (Whole-Body Control) Extending the control target from the end-effector of a robotic arm to the entire body, including the torso, legs, and head, commonly used in humanoid robots
SLAM Simultaneous Localization and Mapping, allowing the robot to know 'where it is and what the surroundings look like'

(1.4) Common Questions for Newcomers

Expand: The 6 Most Frequently Asked Questions Before Getting Started

Q: What prerequisites are needed?
A: Knowing how to use Python, understanding the basic concepts of deep learning (being able to train a classification network), and being able to set up the environment on Linux. Robotics and control theory are not prerequisites, you can learn them as needed when encountering problems, refer to the beginning of the Control Chapter's 'What level should you learn to'.

Q:没有真机器人,还能做具身智能吗?
A:能,而且大多数人都是这么开始的。仿真器 + 公开数据集足以支撑从入门到发论文的完整链路,本指南第 2 章就是纯仿真的。真机的价值在于暴露仿真掩盖掉的问题(标定误差、延迟、接触不确定性),但它应该是第二步而不是第一步。如果确实想碰真机,LeRobot 生态的 SO-101 主从臂是目前最低成本的入口。

Q:需要什么显卡?
A:训练 ACT、Diffusion Policy 这类轻量策略,16GB 显存够用(本指南第 2 章的教程约需 12GB)。微调 π0、GR00T 这类 VLA 基座通常需要 40GB 以上,或使用 LoRA 等参数高效方法。仅做推理部署时要求低得多,部分模型可在消费级显卡上运行。

Q:该先做操作(Manipulation)还是运动(Locomotion)?
A:看你更接近哪边。有 CV / 深度学习背景的人在操作方向上手更快,因为它本质是「带动作输出的监督学习」;有控制、自动化背景的人往往在运动控制方向更有优势,因为它高度依赖仿真中的强化学习与 Sim2Real。两者的技术栈差异比想象中大,建议先选一个做深。

Q:具身智能和传统机器人学是什么关系?
A:不是替代关系。传统机器人学提供几何、动力学与控制的底座,保证系统稳定可控;具身智能在其上用学习的方法解决「感知复杂世界、泛化到没见过的任务」这类传统方法难以手工设计的部分。真实系统几乎都是二者的混合。

Q:现在入场是不是太晚了?
A:相比语言模型,具身智能远没有收敛——数据、评测、本体三个层面都还没有公认答案。2026 年公认的瓶颈已经从模型架构转向数据配方与评测协议(见算法篇 (5.4)),这意味着不依赖超大算力的研究空间依然很多。

(1.5) About Us —— 关于我们

我们是一个由具身智能初学者组成的团队,希望以自己的学习经验为后来者提供帮助,加快具身智能的普及。欢迎更多朋友加入项目,也欢迎交友与学术合作。有任何问题可联系邮箱 [email protected]。

Contributors

⚒️ (2) 动手学习具身智能操作

目标:以 RoboTwin 2.0 为例,完整走通一次操作策略的「生命周期」——读论文建立认知、装环境、拿数据、训练 ACT、跑评测。这条链路本身是通用的,换成别的平台也是同样几步。

前置条件:一块显存不低于 16GB 的显卡(ACT 训练约需 12GB)。官方建议数据采集与策略评测避开 A / H / V 系列显卡,详见 Common Issue。

(2.1) 为什么选择这个教程

具身智能操作是一个复杂的系统问题,可以拆成三个核心环节:

  • 数据从哪来:常见来源包括真机采集、人类视频、仿真合成与世界模型合成,各有短板——真机采集成本高、人类视频信息含量低、仿真合成面临 Sim2Real Gap 与 Scale-up 难题、世界模型合成存在幻觉;
  • 策略怎么设计:网络架构的选择直接影响模型表现、收敛效果与推理速度;
  • 怎么评测性能:没有科学的评测,就无法判断模型好坏,也难以推动技术进步。

面对以上问题,RoboTwin 2.0(ICML 2026)提供了一个很好的学习平台。它基于易配置的 SAPIEN 仿真平台开发,提供 50 个双臂任务的自动化数据合成与统一评测系统,并已开源 10 万条以上预采集轨迹——这意味着新手可以跳过最耗时的数据采集环节,直接从训练开始。

需要注意的是,策略侧的代码现在放在 XPolicyLab 这个子模块里,训练和评测脚本都从那里调用,所以下面的步骤会用到它。它把不同策略的训练与评测收敛到了同一套接口,对学习者的实际好处是:走完一遍 ACT 之后,想换成 π0、RDT-1B 这类 VLA 试试,主要改的是策略名和配置文件,不用从头再理解一套工程。

过程中建议多看合成数据与评测回放视频,以建立对数据分布和策略失败模式的直觉。

(2.2) 学习流程

RoboTwin 2.0:代码|主页|文档|论文

XPolicyLab(策略侧代码):代码|文档

任务与榜单:50 个双臂任务说明|Leaderboard

跑完之后可以看看:RoboDojo(仿真 + 真机统一评测)|RMBench(记忆依赖操作)

(2.2.1) 了解 RoboTwin 2.0 做了什么(约 1 天)

阅读 RoboTwin 2.0 论文,了解仿真数据合成的方案,深入理解合成一条机器人数据需要哪些信息、机器人可以完成什么任务,并了解 ALOHA 硬件。

(2.2.2) 安装平台(约 0.5 天)

按安装文档配置环境,整个过程约 20 分钟。注意 XPolicyLab 现在是 RoboTwin 的 Git 子模块,克隆时必须递归拉取,否则后续训练与评测会缺文件:

# 全新克隆
git clone --recurse-submodules https://github.com/RoboTwin-Platform/RoboTwin.git
cd RoboTwin

# 若此前已克隆过,补拉子模块即可
git submodule update --init --recursive XPolicyLab

(2.2.3) 准备数据(约 0.5 天)

官方已开源 10 万条以上轨迹,推荐直接下载,可以省掉数小时的采集时间。下面只取本教程用到的 beat_block_hammer 任务:

# 只下载指定任务;不传参数则下载全部任务
bash scripts/download_xpolicylab_data.sh beat_block_hammer

数据会落在 data/demo_clean/<task_name>/aloha_agilex/data/。

只有当你需要自定义任务配置、域随机化或更换机器人本体时,才需要自己采集。采集脚本会先搜索能成功完成任务的随机种子,再回放种子录制轨迹:

bash collect_data.sh ${task_name} ${task_config} ${gpu_id}

# Clean Data Example
bash collect_data.sh beat_block_hammer demo_clean 0

# Randomized Data Example
bash collect_data.sh beat_block_hammer demo_randomized 0

自采数据落在 data/<task_config>/<task_name>/<embodiment>/data/,已经是 XPolicyLab 轨迹格式,不需要额外转换。想理解 demo_clean 与 demo_randomized 的差别,可读域随机化文档。

⚠️ 读取 HDF5 里的图像时只能用 XPolicyLab.utils.process_data.decode_image_bit。自己写 cv2.imdecode 或 PIL 解码会因为历史数据版本的布局差异而静默颠倒 RGB 通道,这是最容易踩、也最难排查的坑。

(2.2.4) 训练 ACT 策略(约 1 天)

ACT 是非常经典的操作策略算法,适合作为第一个复现对象,训练大约需要 12GB 显存。策略适配器位于 XPolicyLab/policy/ACT/,XPolicyLab 里所有策略都遵循同一套生命周期脚本:

cd XPolicyLab/policy/ACT
bash install.sh                                  # 安装策略侧运行环境
bash process_data.sh <bench_name> <ckpt_name> <env_cfg_type> <action_type>
bash train.sh <bench_name> <ckpt_name> <env_cfg_type> <action_type> <seed> <gpu_id>

同一套参数命名会贯穿数据处理、训练与评测,中途不需要改名:bench_name 取 RoboTwin;ckpt_name 是本次训练的简称,例如 act_demo;env_cfg_type 用 arx_x5(对应 RoboTwin 默认的 aloha-agilex 布局);action_type 一般取 joint 或 ee。权重会落在 checkpoints/<bench_name>-<ckpt_name>-<env_cfg_type>-<action_type>-<seed>/。具体参数与显存要求以 ACT 适配器 README 为准。

(2.2.5) 评测策略并对照榜单(约 1 天)

RoboTwin 的所有评测都统一走 scripts/eval_policy.sh。调度器会为每个任务各起一个策略服务端和一个仿真器,任务列表与 GPU 配置写在 env_cfg/eval/all_tasks.yml——只想评单个任务的话,把 tasks 裁成一条即可:

bash scripts/eval_policy.sh multitask \
  --config env_cfg/eval/all_tasks.yml \
  --policy-name ACT \
  --ckpt-name <checkpoint> \
  --env-cfg-type arx_x5 \
  --policy-conda-env <policy_env> \
  --eval-env-conda-env <robotwin_env> \
  --action-type joint

加 --dry-run 可以只校验调度计划而不真正启动,结果默认写到 eval_result/multitask/。如果本机显卡不够,还可以把策略服务端放在远程机器、仿真器留在本地,用 --enable-remote 搭配 --policy-server-ip / --policy-server-port 连接,详见 XPolicyLab 文档。

Run after completing the task and compare your success rate with the Leaderboard official results on ACT (about 56% of demo_clean). A significant gap usually indicates misalignment in data volume, training rounds, or action_type.

You have now completed the entire lifecycle of an operational strategy. The most cost-effective next step is to run the same strategy again with a different one: replace --policy-name with any of Pi_0, RDT_1B, GR00T_N17, or X_VLA, and the rest of the process remains the same. This allows you to intuitively feel the differences in various architectures on the same task. The complete strategy list is available in the XPolicyLab documentation.

📄 (3) Useful Info —— Resources for building understanding

This chapter is used for quickly establishing an overall understanding of the field of embodied intelligence, suitable to be read before systematic learning of algorithms, engineering, or hardware, to understand the technical landscape, community ecosystem, and research trends.

(3.1) Directional and Methodological Resources

  • Basic technical roadmap of embodied intelligence (Yunlong Dong): bilibili
  • Introduction to Stanford Robotics: bilibili
  • Cyber Nachos (biased towards system and engineering thinking): website

(3.2) Communities and自媒体 (High long-term follow-up value)

Chinese (WeChat Official Accounts)
Shi Ma Diary, Lumina Embodied Intelligence, Machine Heart, Xin Zhi Yuan, Quantum Bit, Embodied Intelligence Research Room, Embodied Era, Human Five, Xbot Embodied Knowledge Base, Embodied Intelligence Heart, Autonomous Driving Heart, 3D Vision Workshop, Jiangmen Venture Capital, RLCN Reinforcement Learning Research, CVHub

Chinese (Xiaohongshu Bloggers)
WhynotTV, TianxingChen (Chen Tianxing), Mu Yao_YaoMarkMu, Xu Huazhe Harry, Zhou Boyu, Gao Fei, Li Hengyang, Zhu Zheng, Ding Yan, YY Shuo, Mango-Man, RHOSLab #PI-Li Yonglu, Zhenghe Shiyi, Xinyan Ren Yongliang, York Yang-Dyna Robotics, Zhe Lun Class President, Wu Yi, Ding Wenchao, Chen Siheng, Han Xiaoguang, Liang Junwei

There was once an AI scholar Ask Me Anything event on Xiaohongshu, where several professors in the fields of embodied intelligence and robotics answered questions, and the highlighted recording can serve as a clue to find accounts worth following.

Communities and wiki

  • Simulately (community-maintained simulation wiki, the first stop when selecting a simulator): website
  • Jushen Study Society: website

English newsletter / media / podcast

  • Import AI (Jack Clark, biased towards policy and industry perspective): newsletter
  • The Batch (DeepLearning.AI, strong in overview and low threshold): newsletter
  • Ahead of AI (Sebastian Raschka, training and implementation details): newsletter
  • Humanoid Daily (daily updates on humanoid robotics industry): website|Machine Dawn: website
  • IEEE Spectrum Robotics: website|The Robot Report (biased towards commercialization and supply chain): website
  • TWIML AI Podcast: podcast|The Robot Brains (hosted by Pieter Abbeel): podcast

(3.3) Laboratories and Academic Ecosystem

(3.4) High-Quality Conferences and Journals (Focus on when searching for papers)

Robotics: Science Robotics, TRO, IJRR, JFR, RSS, RAL, IROS, ICRA, CoRL

Computer Vision: CVPR, ICCV, ECCV|Machine Learning: NeurIPS, ICML, ICLR|AI and NLP: AAAI, ACL

(3.5) Paper Lists (Focus on for tracking research progress and topic research)

  • Awesome Humanoid Robot Learning (Yanjie Ze): repo
  • Paper Reading List (DeepTimber Community): repo
  • Paper List (Yanjie Ze): repo
  • RoboScholar / Embodied AI Paper List (Tianxing Chen): repo
  • Awesome LLM Robotics: repo
  • Awesome Video Robotic Papers: repo
  • Awesome Embodied Robotics and Agent: repo
  • awesome-embodied-vla / va / vln: repo
  • Awesome Affordance Learning: repo
  • Embodied AI Paper TopConf: repo
  • Awesome RL-VLA for Robotic Manipulation (Haoyuan Deng): repo
  • Awesome Efficient-VLA for Robotic Manipulation (Weifan Guan): repo
  • Awesome Embodied Data: project|repo|arXiv

(3.6) Annual Trend Summaries

  • Data issues in 1,228 VLA papers (June 2026, empirical analysis: VLA paper count increased by 5 times year-on-year): website
  • Survey of VLA datasets / benchmarks / data engines (April 2026): arXiv|repo
  • State of Robot Learning (December 2025): website
  • Xu Huazhe - Embodied Intelligence: Looking Back on 2025: website
  • Lin Tianwei - Embodied VLA in 2025: The Distance from Demo to Generalization: website

🍎 (4) Algorithm —— Algorithm Chapter

Full content: topics/algorithm.md

This chapter connects the most commonly used 'algorithm capability stack' in embodied intelligence from bottom to top: the foundation is engineering tools and geometry, calibration, and control, which determine whether the system can run stably; the middle layer is vision and multimodal representation (2D/3D/4D, prompting, affordance), responsible for compressing the complex world into a middle representation that is generalizable, aligned, and usable by the policy; the upper layer is learning and decision-making (RL/IL, VLA, LLM + Planner, fast-slow systems), transforming perception and task goals into executable actions, and gradually moving towards longer-term, more general, and more deployable system forms.

🏋️‍♂️ (5) Infrastructure —— Software Infrastructure

Complete content:topics/infrastructure.md

This chapter focuses not on 「a specific model」, but on the software infrastructure (Infrastructure) that supports embodied intelligence research and system implementation. Simulators determine what kind of world you can build, benchmark sets determine how you compare method effectiveness, and datasets determine what kind of behavior distribution the model ultimately learns. Together, they form the part most easily overlooked but most influential on performance and reproducibility in embodied intelligence.

🎮 (6) Control —— Control

Complete content:topics/control.md

This chapter is not about getting you to 「run a model immediately」, but to provide stability, interpretability, and engineering foundation for embodied intelligence systems. Control theory ensures the system does not crash under high frequency, robotics provides geometric and dynamic constraints, SLAM and state estimation let the robot 「know where it is」, and ROS and engineering libraries turn theory into a reproducible system.

🦾 (7) Hardware —— Hardware

Complete content:topics/hardware.md

Embodied intelligence hardware covers multiple technology stacks: embedded software and hardware, mechanical design, robot system integration, and sensors, etc. They have a very broad knowledge base, but their common goal is one: to turn 「algorithms」 into stable and reproducible systems in the real world. Regarding hardware learning, the most effective way is almost always to start with practice —— first build a minimal system that can run, then gradually expand complexity and reliability.

🤝 Contributing —— Contributing

This project is maintained by the community. We welcome any form of participation:

  • Supplementary Materials: Add them in the topics/*.md section of the corresponding chapter in the existing format, and try to add a sentence explaining what problem this material solves;
  • Correction of Content: If you find a broken link, outdated conclusion, or incorrect statement, please feel free to submit a PR or open an Issue;
  • Improvement of Structure: If you have ideas about chapter organization and learning routes, please discuss them in an Issue.

Before submitting a PR, please confirm: the entry is placed in the most semantically relevant chapter, the link is accessible, and the format is consistent with adjacent entries.

👍 Citation —— Citation

If this repository has been helpful to you, please cite it:

@misc{embodiedaiguide2025,
  title = {Embodied-AI-Guide},
  author = {Embodied-AI-Guide-Contributors, Lumina-Embodied-AI-Community, Tianxing Chen},
  month = {January},
  year = {2025},
  url = {https://github.com/TianxingChen/Embodied-AI-Guide},
}

🏷️ License —— License

This project uses the Non-Commercial Use license:

  • Allowed: Personal learning, academic research, and other non-commercial purposes;
  • Prohibited: Any form of commercial use, including but not limited to internal use in companies/businesses, integration into paid products or services, or use for any profit-making purpose.

Please check the LICENSE file in the repository for details. If you need commercial authorization (e.g., for use in company products or commercial projects), please contact the project leader: [email protected].

⭐️ Star History —— Star History

Star History Chart

🤝 Sponsors —— Sponsors

Thank you to Wujie Zhihang, Chao Wei Dong Li, HKU MMLab, Dizhu Robot, for their support of this project.

Sponsors

Source:Chinese robotics — Dataset and collection discovery · github.com

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).