Skip to content
Source
Chinese robotics — Hardware and sensing·· 2 hours agoSignalEditorial score85

Hand-Eye-Brain Coordination: A Full Guide to Industrial and Service Robot Integration

手眼脑协同:工业与服务机器人产业化落地全解析 - 社区

Summary

This article explains the 'hand-eye-brain' coordination framework for industrial and service robots, detailing the roles of perception, decision-making, and execution. It covers challenges like sensor calibration, force control, and system synchronization, and offers practical advice on simulation, testing, and deployment in real-world scenarios.

Source: Chinese robotics — Hardware and sensing · Read original article ↗

Article text · Machine translation into English

On this page

1. From the perspective of 'Hand-Eye-Brain' coordination: Why this direction is worth exploring

The first time I heard the term 'Hand-Eye-Brain' coordination was during a post-mortem meeting for an industrial robot integration project. At that time, a six-axis robot was picking up randomly stacked parts on the production line, but its success rate was stuck at 85% and couldn't be improved. The vision system provided coordinates, the robotic arm was in place, but the grasping posture was always slightly off. Later, a teacher who specialized in robot theory pointed out: 'You're fighting separately with your 'eye' and 'hand', missing a 'brain' to make unified decisions.' This sentence has stayed with me for many years and basically summarizes the core research focus of Professor Qiao Hong's team over the years.

What is 'Hand-Eye-Brain' coordination? Breaking it down, it consists of three parts: Hand is the execution mechanism, which refers to physical entities such as robotic arms, end-effectors, and dexterous hands; Eye is the perception system, including 2D/3D vision, force sensing, tactile sensing, and other multimodal sensors; Brain is the decision-making and control center, responsible for converting perceptual information into action commands and making real-time adjustments during execution. Individually, these three components have already been quite mature in both industry and academia, but the real challenge lies in making them work together like a human's hand, eye, and brain - seeing what's there, thinking about what to do, and acting accordingly, forming a closed loop.

Why is this direction particularly worth paying attention to now? Because industrial robots and service robots are currently at a critical point of transition from 'teaching and repeating' to 'autonomous decision-making'. Traditional industrial robots rely on pre-programmed trajectories, with fixed workpiece positions, stable lighting conditions, and clear timing requirements. In such scenarios, the coordination between 'hand' and 'eye' can be solved through calibration and offline programming. However, once the scenario becomes random grasping, flexible assembly, human-robot collaboration, or service robots operating in unstructured environments like homes, hospitals, or malls, pre-set trajectories become completely ineffective. At this point, a 'brain' is needed to process uncertainty in real-time.

Professor Qiao Hong's research path precisely hits this critical point of transition. She started with robot theory and control, then focused on 'Hand-Eye-Brain' coordination, and later promoted industrialization. This line of research actually reflects the inevitable logic of the robotics discipline moving from theory to application. For those developing robots, understanding this framework is far more important than simply learning a particular algorithm or adjusting a parameter - because it determines the top-level design approach when building a system.

I've seen many teams start by stacking hardware: buying the best depth cameras, using the most expensive collaborative arms, and equipping the most advanced force-controlled end-effectors. However, after system integration, they found that the modules didn't communicate well, the timing wasn't aligned, and the decision logic conflicted. The problem lies in not designing the overall architecture from the perspective of 'Hand-Eye-Brain' coordination. Therefore, this article aims to explain the 'Hand-Eye-Brain' coordination from theory to implementation, combining the research trajectory and industrialization practices of Professor Qiao Hong's team, and also to discuss the pitfalls on the path to industrialization of industrial and service robots.

2. Dissecting the Core Architecture of 'Hand-Eye-Brain' Coordination

2.1 'Brain' Layer: Algorithm Stack from Perception to Decision

The 'brain' is the central system of the entire setup, with its core task being to fuse multimodal perceptual information into a consistent understanding of the environment, and then generate action sequences based on that. In the research of Professor Qiao Hong's team, the 'brain' layer typically includes several key modules: State Estimation, Task Planning, Motion Generation, and Online Adjustment.

State estimation addresses the question of 'where am I and what's the situation around me'. In industrial scenarios, this includes workpiece pose estimation, robot joint states, and external force feedback; in service scenarios, it also includes human detection, semantic map building, and voice command understanding. The challenge here lies in the fusion of multi-source heterogeneous data - vision provides pixel coordinates, force sensing provides Newton values, encoders provide angles, and the question is how to unify them into a single state space, which is the first hurdle in algorithm design.

Task planning is responsible for breaking down high-level goals into executable subtask sequences. For example, the instruction 'put the part on the table into the box' needs to be broken down into steps like 'locate the part → plan the grasping posture → move to the box → adjust the posture → place'. Traditional methods use state machines or behavior trees, but more and more learning-based methods are being used, especially for complex operations.

Motion generation is the process of converting subtasks into specific joint trajectories or end-effector trajectories. This involves classic problems such as kinematics solving, trajectory optimization, and obstacle avoidance planning. Professor Qiao Hong's team has made significant contributions in motion generation, particularly in inverse kinematics solving for redundant robotic arms and dynamic obstacle avoidance.

Online adjustment is the part of the 'brain' layer that best demonstrates intelligence. During execution, if the vision detects that the workpiece position has shifted, or if the force feedback indicates an abnormal contact force, the 'brain' needs to adjust the action in real-time. This requires algorithms with sufficient response speed and robustness.

Practical insights: Many teams make a common mistake when designing the 'brain' layer - splitting planning and control into two independent modules. Planning outputs a fixed trajectory, and control only tracks it. This architecture works fine in structured environments, but once the environment changes dynamically, planning can't recalculate in time, and control can't adjust autonomously, causing the system to freeze. A better approach is to use a layered architecture: the high-level planner provides a rough path, and the low-level control combines real-time perception for local adjustments, with high-frequency communication between the two layers.

2.2 'Eye' Layer: Multimodal Perception and Calibration

'Eye' is the source of information for the system, and its accuracy and robustness directly determine the system's upper limit. In industrial robot scenarios, 'eye' typically refers to the vision system, including 2D cameras, 3D cameras, and laser profilometers; in service robot scenarios, it also includes depth cameras, ultrasonic sensors, and infrared sensors.

The core tasks of the vision system are three: Target Detection and Recognition, Pose Estimation, and Scene Understanding. Target detection addresses the question of 'what is it', pose estimation addresses 'where it is and what posture it has', and scene understanding addresses 'what is the surrounding environment like'. These three tasks have different focuses in industrial and service scenarios: industrial scenarios focus more on the accuracy and speed of pose estimation, while service scenarios focus more on the semantic richness of scene understanding.

Calibration is the most fundamental and easiest-to-go-wrong part of the 'eye' layer. Hand-Eye Calibration solves the transformation relationship between the camera coordinate system and the robot's end-effector coordinate system. Calibration accuracy directly affects grasping accuracy, and calibration errors are often nonlinear, showing different behaviors in different working areas.

I've seen a case: a production line using Eye-in-Hand configuration, with the camera mounted on the robotic arm's end-effector, had a grasping accuracy of 0.1mm in the central area after calibration, but the error increased to over 0.5mm at the edge of the workspace. After investigation, it was found that the calibration board coverage was insufficient, and the distortion in the edge area wasn't fully modeled. After re-calibration with a calibration board covering the entire workspace, the problem was resolved.

Notes: Hand-Eye Calibration is not a one-time task. Over time, joint clearance and link thermal deformation can cause calibration parameter drift. It's recommended to re-validate the calibration after line changeover or after continuous operation for 24 hours. For high-precision applications, online calibration schemes can be considered, using known feature points to real-time adjust the transformation matrix.

2.3 'Hand' Layer: Execution Mechanism and Force Control Strategy

'Hand' is the final part that interacts with the physical world, and its performance determines what operations the system can perform. Common end-effectors for industrial robots include two-finger grippers, three-finger grippers, vacuum suction cups, and electromagnetic suction cups; service robots use more dexterous hands and soft hands.

The choice of execution mechanism should match the task requirements. For grasping regular metal parts, a two-finger gripper with force control is sufficient; for grasping fragile items, a soft hand or a gripper with force feedback is needed; for assembly tasks, a six-axis force sensor for compliant control may be required.

Force control strategy is the core of the 'hand' layer. Position control only ensures reaching the specified position, while force control manages the size and direction of contact force. In tasks like assembly, grinding, and polishing, force control is essential. Common force control methods include impedance control, admittance control, and hybrid force/position control.

Professor Qiao Hong's team has done a lot of work in force control, especially adaptive force control in uncertain environments. Traditional impedance control requires pre-set impedance parameters, but the stiffness of the contact environment is often unknown in reality. If the parameters are set incorrectly, the system may either be too rigid and cause collisions or too soft and result in weak operation. Adaptive impedance control can adjust parameters in real-time based on contact force feedback, adapting to different contact environments.

2.4 Data Flow and Timing Between the Three Layers

The key to 'Hand-Eye-Brain' coordination lies in the data flow and timing coordination between the three layers. A typical coordination process is as follows:

  1. 'Eye' collects environmental data, performs target detection and pose estimation, and outputs target pose and confidence level
  2. 'Brain' receives perceptual results, combines them with the task goal, and generates grasping posture and motion trajectory
  3. 'Hand' executes the motion command, while force sensors provide real-time contact state feedback
  4. 'Brain' adjusts the trajectory based on force feedback and visual tracking results in real-time
  5. After grasping is completed, 'Eye' verifies the grasping result, and 'Brain' decides the next action

In this process, timing is crucial. Visual processing typically takes tens to hundreds of milliseconds, motion planning takes several to tens of milliseconds, and control cycles require a period of 1ms or less. If visual processing is too slow, the 'brain' will receive outdated information, leading to decision errors. Therefore, in actual systems, visual processing often uses ROI cropping, downsampling, and parallel computing to speed things up.

Practical insights: When building a 'Hand-Eye-Brain' system, it's recommended to first validate the algorithm logic in a simulation environment before moving to physical hardware. Gazebo, PyBullet, and MuJoCo are commonly used simulation platforms. In simulation, you can quickly iterate algorithms without worrying about hardware damage. However, the transition from simulation to real-world application (Sim-to-Real) is a major pitfall. The ideal sensor models in simulation and the noise characteristics of real sensors differ significantly. It's recommended to actively add noise and delay in simulation to improve algorithm robustness.

3. Industrial Robot Industrialization: The Path from Theory to Production Line

3.1 Core Needs and Pain Points in Industrial Scenarios

The core driver for industrial robot industrialization is cost reduction and efficiency improvement. Replacing human labor with robots on a production line typically requires an investment payback period of 1-2 years. This means the robot system must be sufficiently cheap, reliable, and easy to maintain.

The main pain points in the current industrial robot deployment are concentrated in several areas:

Insufficient flexibility. Traditional industrial robots are suitable for mass production with few product varieties, requiring re-teaching or programming during changeover, resulting in long downtime. However, modern manufacturing is increasingly moving towards small-batch, multi-variety production, requiring robots to quickly adapt to new tasks.

Weak perception capabilities. Most production line robots are still 'blind grabbing', relying on fixtures to ensure consistent workpiece positions. Once there is a deviation in the incoming workpieces, a visual system needs to be added, but the integration and debugging of the visual system are very complex.

High programming threshold. Industrial robot programming usually requires professionals, and frontline operators find it difficult to master. This leads to dependency on equipment suppliers for production line adjustments, resulting in slow response times.

Long payback period. A complete robot workstation, including the robot itself, end effector, visual system, safety fence, and control system, can cost tens to hundreds of thousands of dollars. For small and medium-sized enterprises, the threshold is relatively high.

The team led by Qiao Hong has promoted the 'hand-eye-brain' collaborative technology, which directly addresses these pain points. By enhancing perception and decision-making capabilities, the flexibility of robots is improved; by simplifying the programming interface, the usage threshold is lowered; and through modular design, the system cost is reduced.

3.2 Typical Application Scenario Breakdown

Scattered Picking (Bin Picking) is the most typical industrial application of the 'hand-eye-brain' collaboration. The workpieces are scattered and stacked in a container, and the robot needs to identify the position and orientation of each workpiece, plan a collision-free grasping path, and then execute the grasping. This scenario has very high requirements for 'eyes': workpieces may be occluded, have reflective surfaces, or complex geometries, all of which can affect the accuracy of pose estimation. The requirements for 'brain' are also high: it needs to handle grasping priorities, collision detection, and path planning issues.

Flexible Assembly is another important scenario. Assembly tasks usually require force control because the tolerance between parts is very small, and pure position control can easily cause jamming or damage to the parts. Force-controlled assembly requires 'hands' with six-dimensional force sensors, 'brain' with force control algorithms, and 'eyes' with visual guidance. The team led by Qiao Hong has conducted considerable research in tasks such as shaft-hole assembly and gear meshing.

Human-Robot Collaboration is a recent hot topic. Traditional industrial robots require safety fences for isolation, while collaborative robots can work in the same space. This requires robots to perceive the position and movements of people and adjust their behavior in real time. This places higher demands on 'eyes' for human detection and 'brain' for safe decision-making.

3.3 Key Elements for Industrialization Implementation

From theory to production lines, there are many engineering challenges in between. I have summarized several key elements:

Reliability over Advancedness. Production lines are not laboratories, and they cannot tolerate frequent failures. An algorithm with 99% accuracy in a paper is very impressive, but 1% failure rate on a production line means dozens of stoppages per day. Therefore, during industrialization, mature and stable solutions are often chosen instead of the most advanced algorithms.

Standardized Interfaces. Interfaces between the robot body, visual system, end effector, and PLC should be standardized to enable quick integration and replacement. Many domestic robot manufacturers are promoting their own interface standards, but the industry as a whole is still fragmented.

Debugging Toolchain. Production line debugging requires intuitive tools to quickly identify problems. For example, calibration tools for visual systems, tools to track the success rate of grasping, and tools to record abnormal situations. These tools often have a greater impact on the efficiency of implementation than the algorithms themselves.

Cost Control. Industrial customers are price-sensitive, and the system cost should be controlled within a reasonable range. Sometimes, a 2D camera combined with structured light can solve the problem, and there's no need to use a 3D camera. Similarly, if a traditional vision algorithm can solve the problem, there's no need to use deep learning.

Practical Insights: When implementing in industrial scenarios, it is recommended to start with a small-scale pilot, selecting a relatively simple workstation to verify the stability and economic viability of the entire system. After a successful pilot, gradually expand the implementation to more workstations. Avoid starting with full-line transformation, as the risk is too high. Additionally, communicate frequently with production line operators, as they understand the actual conditions and anomalies in the production environment, which are very valuable for system design.

4. Industrialization of Service Robots: Challenges and Responses in Unstructured Environments

4.1 Essential Differences Between Service and Industrial Scenarios

The biggest difference between service robots and industrial robots is the degree of unstructured environment. In industrial scenarios, the positions of workpieces, lighting conditions, and operational procedures are controllable; in service scenarios, the positions of objects are random, lighting changes are significant, and people move frequently, requiring robots to handle a large amount of uncertainty.

This difference leads to different technical approaches. Industrial robots can rely on precise calibration and pre-set trajectories, while service robots must depend on real-time perception and online decision-making. The 'brain' of industrial robots can plan most actions offline, while the 'brain' of service robots must handle various unexpected situations online.

Another difference is interaction objects. Industrial robots mainly interact with workpieces, while service robots must interact with people. This requires service robots to have natural interaction capabilities, including speech recognition, gesture understanding, and face tracking. The work of the team led by Qiao Hong in service robots also involves human-robot interaction and emotional computing.

4.2 Special Requirements for 'Hand-Eye-Brain' in Service Robots

The 'eyes' of service robots need stronger scene understanding capabilities. In industrial scenarios, the visual system only needs to identify specific workpieces; in service scenarios, robots need to understand the semantics of the entire scene, such as 'this is a living room, there are clothes on the sofa, and a cup on the coffee table.' This requires the integration of deep learning, semantic segmentation, and scene graphs.

The 'hands' of service robots need higher compliance and safety. Industrial robots can grab forcefully, but service robots handle items like cups, clothes, and food, which are fragile, requiring force control and compliant operations. Soft hands and adaptive grippers are more commonly used in service robots.

The 'brain' of service robots needs stronger task understanding and planning capabilities. User instructions for service robots are often vague, such as 'clean up the room.' The robot needs to understand the meaning of the instruction, decompose it into specific tasks, and handle various unexpected situations during execution.

4.3 Analysis of Typical Service Robot Applications

Home Service Robots are one of the hottest directions currently, but also the most challenging. The home environment is highly unstructured, with a wide variety of items, vague user instructions, and high safety requirements. Currently, the products that have been deployed are mainly focused on single tasks such as sweeping and window cleaning, and general-purpose home service robots still have a long way to go.

Medical Rehabilitation Robots are another important direction. Rehabilitation robots require precise force control and motion control, while also ensuring patient safety. The team led by Qiao Hong has conducted related work in rehabilitation robots, involving human-robot interaction, intent recognition, and adaptive control.

Logistics Service Robots have been widely deployed in warehousing and logistics scenarios. These scenarios are relatively structured, and robots mainly perform tasks such as transportation and sorting. AGV/AMR combined with robotic arms form a 'hand-eye-brain' collaborative logistics robot system.

Education and Entertainment Robots have relatively lower requirements for 'hand-eye-brain,' but higher requirements for interactive experience. These robots typically use simple visual and voice interaction, with actions mainly pre-set.

Notes: When deploying service robots, safety is the top priority. Robots that coexist with humans must have multiple safety mechanisms: hardware emergency stop, software speed limit, force sensing collision detection, and visual human detection. Any failure in any of these layers can lead to a safety accident. It is recommended to refer to safety standards such as ISO 13482 for design.

5. Common Issues and Troubleshooting Techniques

5.1 Troubleshooting Visual-Related Issues

Problem: Low grasping success rate and unstable visual detection

Troubleshooting Approach: First, confirm whether the issue is with detection or pose estimation. If the detection box is jittering, it may be due to changes in lighting conditions or improper camera exposure settings; if the pose estimation deviation is large, it may be due to inaccurate calibration or poor point cloud quality.

Solutions: Fix the lighting conditions and use active lighting; adjust the camera exposure and gain to avoid overexposure or underexposure; recalibrate and ensure the calibration board covers the working space; for reflective workpieces, use a polarizing filter or adjust the shooting angle.

Problem: Insufficient accuracy after hand-eye calibration

Troubleshooting Approach: Check the calibration board accuracy, the distribution of calibration points, and the robot's repeatability accuracy. Calibration points should be evenly distributed across the working space, avoiding concentration in a single area. If the robot's repeatability accuracy is poor, the calibration will be ineffective regardless of its accuracy.

Solutions: Use a high-precision calibration board; increase the number of calibration points to cover the entire working space; check the joint clearance and gear backlash of the robot; for high-precision applications, consider online calibration.

5.2 Troubleshooting Motion Planning and Control Issues

Problem: Mechanical arm vibration during movement

Troubleshooting Approach: Check if the trajectory is smooth and if the acceleration is abrupt; check if the joint servo parameters are appropriate; check for external interference.

Solutions: Use an S-shaped velocity curve instead of a trapezoidal curve; adjust the servo gain to reduce stiffness or increase damping; check the mechanical structure for looseness.

Problem: Jamming or damage to parts during force-controlled assembly

Troubleshooting Approach: Check if the force control parameters are appropriate, whether the impedance is too high or too low; check if the contact force feedback is accurate; check if the assembly strategy is reasonable.

Solutions: Adjust the impedance parameters to increase compliance; calibrate the force sensor to eliminate zero drift; use a search strategy to first find the hole and then insert it; use a spiral search or vibration strategy to handle jamming.

5.3 Troubleshooting System Integration Issues

Problem: Large communication delay and slow system response

Troubleshooting Approach: Check the communication protocol and network configuration; check the processing time of each module; check for synchronization issues.

Solutions: Use real-time communication protocols such as EtherCAT; optimize the visual algorithm to reduce processing time; use multi-threading or parallel computing; unify the clock source to ensure temporal consistency.

Problem: System accuracy decreases after long-term operation

Troubleshooting Approach: Check the impact of temperature drift; check for mechanical wear; check if the calibration parameters have drifted.

Solution: Add temperature compensation; perform regular maintenance on mechanical structures; regularly recalibrate; for critical applications, use online calibration.

Practical insights: When troubleshooting, it is recommended to start with the simplest possibilities. I have encountered "unstable visual inspection" several times, and in the end, it was found to be due to the camera lens not being tightened or poor contact in the network cable. Check hardware connections and basic settings first, then move on to the algorithm level. Additionally, develop the habit of logging. Record the time, phenomenon, and parameters for each anomaly. With accumulated data, you can quickly identify patterns of recurring issues.

6. Toolchain and Learning Path Recommendations

6.1 Simulation and Development Tools

Do 'hand-eye-brain' collaborative development, simulation environment is essential. Gazebo combined with ROS2 is currently the most mainstream solution, supporting various robot models and sensor simulations. PyBullet is lightweight and suitable for rapid algorithm verification. MuJoCo has high precision in contact simulation, suitable for force control research.

ROS2 is the standard framework for robot development, providing basic functions such as communication, coordinate transformation, and visualization. MoveIt2 is a motion planning library, supporting various planning algorithms. OpenCV and PCL are fundamental libraries for visual processing. PyTorch and TensorFlow are used for deep learning model development.

For industrial robots, each brand has its own programming environment, such as ABB's RobotStudio, KUKA's KUKA.WorkVisual, and FANUC's RoboGuide. These environments typically support offline programming and simulation.

6.2 Learning Path Recommendations

Introductory stage, it is recommended to first learn ROS2 basics, understand concepts such as nodes, topics, services, and actions. Then learn robot kinematics, understand forward and inverse kinematics, Jacobian matrix, and trajectory planning. At the same time, learn computer vision basics, understand camera models, calibration, and feature extraction.

Advanced stage, learn motion planning algorithms, such as RRT, PRM, and optimization-based methods. Learn force control algorithms, such as impedance control and admittance control. Learn the application of deep learning in robotics, such as grasp detection, pose estimation, and reinforcement learning.

Practical stage, find a specific application scenario, and do it from simulation to physical implementation. For example, scattered grasping, shaft-hole assembly, and mobile grasping. During the process, you will encounter various problems, and the process of solving these problems is the best learning.

6.3 Recommended resources

Books: "Introduction to Robotics" is a classic textbook, "Modern Robotics" leans more towards algorithms, and "Probabilistic Robotics" covers SLAM and state estimation. For ROS2, "ROS2 Robotics Development from Beginner to Practice" is a good introductory book. For vision, "Computer Vision: Algorithms and Applications" is quite comprehensive.

Open source projects: MoveIt2, Nav2, OpenVINO, and Detectron2 are all excellent learning resources. There are many open-source implementations on GitHub for robotic grasping and assembly, which can be referenced.

Finally, here's a small tip: when doing "hand-eye-brain" collaborative development, it's recommended to first build a minimum viable system (MVP), using the simplest hardware and algorithms to run through the entire process first, and then gradually replace and optimize each module. This allows for quick validation of the architecture design and avoids excessive investment in details. I've seen too many teams get stuck on a single module, only to find that the overall architecture was flawed, leading to a complete rebuild. The order of first running through and then optimizing is very important.

Source:Chinese robotics — Hardware and sensing · bbs.csdn.net

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).