Focusing on Autism and Other High-Interaction Scenarios, This Company is Defining Robot 'Social' Interaction
聚焦孤独症等强交互场景,这家公司正在定义机器人“社交”
Wuhan Huilike Robotics Co., Ltd. has developed the VLI multimodal embodied interaction model based on research achievements from a team at Huazhong University of Science and Technology. The company focuses on enhancing robots' social capabilities in real-world scenarios. Through an L0 to L5 interaction capability grading framework, the company defines multiple levels of social capabilities ranging from tool execution to non-cooperative strong active interaction. In specialized scenarios such as autism assistance intervention, Huilike combines expert experience with multimodal data to enable robots to achieve the goals of understanding humans, adjusting strategies, and providing continuous service.
Source: RoboSpeak — WeChat · Read original article ↗
Article text · Machine translation into English
“
Let robots enter human society.
01.
Why is multimodal 'social' interaction most important for robots entering human society?
Can robots chatting be considered as 'social' interaction?
When embodied robots enter real service scenes, a wave of the hand, a look seeking help, or even hesitation in tone may change the robot's next response. It needs to combine language, expressions, and actions to determine when to speak, when to wait, and adjust interactions according to human feedback.
Based on
the team from Huazhong University of Science and Technology, which has been established as Wuhan Huilike Robotics Co., Ltd.
the company is focusing on this direction. The company is centered around its self-developed VLI multimodal embodied interaction model, exploring how robots can understand human states, adjust social strategies, and continuously complete professional services in real scenarios.
Figure 1: Huilike Robotics Logo
02.
From L0 to L5, defining the boundaries of robot interaction capabilities
Huilike defines the social interaction capabilities of robots as L0 to L5 levels, based on the tasks robots are responsible for and the way human intervention occurs. The framework also examines system proactiveness, multimodal coordination, individual memory, and scene complexity, clearly defining the interaction tasks required at each level, providing a common scale for product development, scenario selection, and effect evaluation.
L0: Follow instructions → L1: Perceive → L2: Understand context → L3: Face individual → L4: Face group → L5: Non-cooperative interaction
Figure 2: Huilike Embodied Robot Human-Robot Interaction Levels (L0–L5)
L0: Tool execution
Interaction is initiated and controlled by humans, with robots executing actions based on clear instructions or fixed scripts. Task goals, execution order, and exception handling are mainly decided by humans, and the system has not yet developed the ability to adjust interactions based on the state of the object. Remote control, button triggering, and fixed action execution are typical forms of this level.
L1: Passive multimodal response
Robots can respond after humans initiate or clearly input, combining language, vision, and actions for expression. Interaction mainly revolves around the current request, with task progression relying on continuous guidance from humans. Further coordination of multiple modalities is needed to support stable goal closure.
L2: Contextual closed-loop interaction
In limited scenarios or clear goals, robots can organize tasks through multi-round interaction, adjusting steps based on multimodal feedback and process memory, and forming recordable and evaluable results. The key is whether the system can maintain the goal, use existing information, handle task branches, and complete the process. Professional personnel set task boundaries, review results, and handle situations beyond the scope.
L3: Individual closed-loop interaction
Robots, on the basis of scenario closed-loop interaction, continuously adjust strategies by combining individual characteristics, historical feedback, and cross-session memory. Information from the previous interaction can influence the next task, and interaction methods update according to the individual. The key for L3 is whether the same object can receive continuous and adapted responses in multiple 'social' interactions, and whether these adjustments have traceable basis.
L4: Interaction with groups
Robots can understand participant relationships, task changes, and on-site constraints in social scenarios, coordinating multi-person or multi-robot interactions. The system needs to adjust interaction objects, rounds, and actions within the scope of the scene authorization. Evaluation focuses on expanding from continuous interaction with a single object to the coordination ability of the group in complex scenarios.
L5: Non-cooperative strong active interaction
The system cannot rely on the other party initiating questions, clearly expressing, or cooperating with a preset process. It needs to actively create opportunities for communication based on the object's state, understand subtle or uncertain feedback, and continuously adjust the timing of intervention, expression methods, and task rhythm. Children's autism assistance intervention is a key scenario that Huilike is exploring at this level, with specific interactions following the goals and boundaries set by professionals.
Judgment at each level should be conducted within clear tasks, objects, and environments, and record task results, multimodal coordination performance, basis for strategy adjustments, and human intervention. The capability level is supported by verifiable interaction processes. Huilike hopes to gradually establish a scale for judging robot social interaction through this grading and evaluation proposition, and continuously improve it in practical applications.
03.
“Goal-Driven” from L5 to L2, the product thinking behind the implementation scenarios
The product path of HuiLiKe starts from a high difficulty interaction problem: when the interaction object cannot actively express needs or continuously participate in the conversation, how can a robot establish and maintain interaction?
Children with autism spectrum disorder (ASD) auxiliary intervention is the key scenario that the company has chosen to address this issue.
Most children with autism have difficulties in expressing their needs, understanding others' expressions and gestures, and participating in a back-and-forth conversation. When facing such children, professionals need to first find what they are interested in, create opportunities for interaction, and then judge the next step based on a single gaze, an action, or a brief response: continue guiding, patiently wait, or try a different way of expression. Most of the time, ordinary people also find it difficult to interact normally with children with autism.
Figure 3: A therapist wearing HuiLiKe's data collection equipment performs rehabilitation on the child
HuiLiKe refers to this type of interaction where the system does not require the other party to actively ask questions, clearly express needs, or follow a preset process as “non-cooperative interaction.” The term “non-cooperative” describes the interaction conditions and does not mean that the child is intentionally refusing to cooperate. It requires the robot to actively find entry points, understand subtle feedback, and continuously adjust strategies according to individual differences, corresponding to the company's long-term goal of L5 non-cooperative strong active interaction.
The value of this scenario also lies in how experts complete the interaction. When do doctors and therapists intervene, based on what feedback they change strategies, and how to keep the conversation going are all professional experiences that the model needs to learn. In the field of autism auxiliary intervention, HuiLiKe is collaborating with authoritative experts such as Professor Hao Yan.
Professor Hao Yan is a Fellow of the Royal Society of Medicine in the UK, a member of the Standing Committee of the World Chinese Association for Prenatal and Postnatal Care, and the Vice Chairman of the Chinese Rehabilitation Association's Committee on Autism Rehabilitation. He has strong professional experience in child development assessment, autism diagnosis and treatment, and rehabilitation guidance.
This collaboration provides HuiLiKe with three aspects of support:
Expert guidance clarifies the research and development direction, professional interaction data supports model learning, and clinical resources connect real application needs.
The team has followed the experts into diagnosis and intervention sites, understanding how doctors and therapists observe children, initiate interactions, judge feedback, and adjust strategies. They have collected multimodal data covering the diagnosis and intervention process, gradually transforming professional experience into tasks and interaction methods that the model can learn and robots can execute.
By leveraging the communication foundation established by the expert team in serving children and families, HuiLiKe can more directly understand actual needs, connect with diagnostic and rehabilitation scenarios, and create conditions for product trials, feedback collection, and continuous validation. The expert collaboration thus spans from need definition, data accumulation, and scenario validation, supporting the company's long-term efforts toward L5 non-cooperative strong active interaction. At the same time, it enables the development of products in scenarios that are ready for delivery, which is HuiLiKe's approach to implementation.
In Wuchang Tanhualin, the company has already implemented a middle school student career planning interaction scenario:
The robot, driven by the VLI specialized brain, engages in conversations around students' interests, concerns, and planning goals, organizes information according to the expert framework, guides expression, and forms planning recommendations, verifying the L2 stage's multi-turn interaction and goal closure capabilities.
In the middle school student career planning scenario in Wuchang Tanhualin, HuiLiKe collaborates with Mr. Wenlong, an educator with 17 years of teaching experience and who has served over a thousand students. Mr. Wenlong has a background in computer science and educational psychology, has been invited to speak at Harvard, Columbia University, the University of Pennsylvania, and several 985 universities in China, and has long focused on the integration of academic, personal development, and career education for teenagers, forming personalized growth plans.
His practice in Beijing is characterized by long-term companionship, combining subject learning, individual understanding, and growth planning, accumulating practical experience in academic improvement and continuous follow-up with students' development. What HuiLiKe values is the method of continuously understanding, guiding, and adjusting around specific students: how to discover interests and concerns, how to break down long-term goals into phased tasks, and how to adjust the plan based on feedback.
Figure 4: HuiLiKe's interaction scenario in the Tanhualin middle school career planning store
Both parties, based on these educational experiences and case accumulations, have transformed the expert methods into a VLI-driven career planning interaction process, allowing the robot to assist students in organizing their own situations, clarifying goals, and forming action recommendations, verifying the L2 stage's goal closure capabilities in Tanhualin. From providing deep private services to a small number of families, the company is moving toward educational普惠 (universal) attempts for more students.
On this basis, the team plans to advance toward L3 individualized strategy coordination through cross-session memory and individual feedback, further expanding toward L4 by enhancing scene understanding and interaction organization capabilities. Each scenario accumulates data and verifies effectiveness, collectively supporting the iteration of multimodal perception, interaction decision-making, and collaborative expression capabilities. Using high-demand scenarios to drive R&D and forming deliveries through phased applications, the long-term research and actual implementation continuously support each other.
04.
The VLI Interaction Specialized Brain Supporting Interaction Capabilities
The VLI Interaction Specialized Brain is the core technology and product of HuiLiKe, carrying the company's R&D on robots understanding humans, actively interacting, and providing continuous services. Based on the team's research accumulation in robot perception and interaction, planning and control, and body systems, HuiLiKe further integrates visual perception, voice interaction, emotional expression, and action generation, combining single-modal and multimodal research results to propose and independently develop the VLI multimodal embodied interaction large model (Vision-Language-Interaction Model), which serves as the model core of the interaction specialized brain, integrating human state understanding, individual memory, interaction strategy decision-making, and collaborative expression into a unified framework.
Around the VLI Interaction Specialized Brain, the robot terminal handles perception and expression, while expert knowledge, interaction data, and scenario feedback support capability training and iteration. Focused on professional needs such as autism auxiliary intervention and career planning, HuiLiKe further integrates expert knowledge, work processes, and interaction methods into the specialized brain, forming professional service capabilities suitable for different tasks, and continuously validates and improves them through practical applications, gradually building an interaction foundation that can be adapted to different robot terminals and expanded to various professional scenarios.
Continuous Emotional Modeling and Biomimetic Collaborative Expression
In emotional interaction research, the team proposes a continuous dynamic emotional representation method, enabling emotional states to evolve synchronously with language generation, providing a foundation for emotional continuity and expression control in ongoing conversations. In terms of physical expression, the team has developed three generations of humanoid robot heads, achieving联动 (synchronized) expression in the eyes, mouth, and facial areas, and established a mapping method between speech and dynamic lip shapes, converting emotions and language into executable facial actions.
These studies respectively form the technical foundation for emotional modeling, voice-driven interaction, and biomimetic expression. HuiLiKe further integrates them into its self-developed VLI multimodal embodied interaction large model, exploring coordinated language, tone, expression, and action under a unified interaction strategy, enabling robots to adjust the content, intensity, and rhythm of expression based on human feedback. In the context of autism assistance intervention, this capability is directed toward specific tasks, providing clear, coherent, and individualized interactive demonstrations under the guidance of professionals, allowing multiple expressions to serve the same interaction goal.
Embodied Action Model and Robot System Accumulation
The team has conducted research such as DeepThinkVLA in the direction of embodied operational models, exploring the generation of robot actions by combining vision, language, and reasoning processes; on the engineering side, they have accumulated experience in robot planning and control, multi-degree-of-freedom execution mechanisms, and whole-machine integration. These achievements provide methods and verification platforms for the transformation of interaction intent into physical actions.
Towards VLI, related research is being integrated with interaction strategies, enabling speech, gaze, head posture, and hand movements to execute collaboratively around the same goal. The team incorporates model outputs, action timing, and terminal operational constraints into integrated verification, pushing algorithm capabilities into real service processes.
Figure 5: HuiLiKe's developed StarBaby Inquiry Assistant entering the outpatient department
Expert Methods and Individual Memory Scenario Transformation
HuiLiKe's collaboration with hospital experts has progressed to the observation of diagnosis, developmental assessment, and intervention interaction processes. The team focuses on what feedback experts use to make judgments, when to change guidance methods, and how to evaluate an interaction, then organizes these experiences into task processes, strategy bases, and evaluation requirements.
On this basis, VLI builds process memory around the current task and explores the use of authorized individual information, historical interactions, and expert feedback for subsequent strategy adjustments. Professional knowledge, interaction experience, and individual changes thus enter the same development process, providing a basis for continuous judgment in long-term services.
Multimodal Data Collection for Real Interaction
The team has already conducted research in the synchronous collection of visual, tactile, and motion states, and has developed head-mounted data collection devices tailored for interaction scenarios, incorporating first-person perspective, facial expressions, speech, and human movements into the data collection plan. The focus is on preserving the temporal relationship between expert expressions and object feedback, enabling the interaction process to be traced back and analyzed.
These data, combined with task processes, expert evaluations, and model outputs, will be used to analyze the applicability conditions of interaction strategies and support system iteration. The team is continuously connecting expert methods, multimodal data, model strategies, and robot execution. This cross-link development and verification capability is a crucial foundation for HuiLiKe to form its own technical characteristics.
05.
From Interaction Framework to Real Service
For industry clients, whether a robot's 'social' capability is valid lies not in single-point demonstrations, but in its ability to be embedded in real service processes, forming deliverable, evaluable, and replicable system capabilities.
HuiLiKe is advancing product development around the 'interaction model + robot terminal + scenario solution': VLI specialized brain provides understanding, memory, and interaction strategy capabilities, robot terminals take on perception and collaborative expression, and professional scenarios define task goals and evaluation requirements. All three are indispensable, and only through synergy can robots transition from 'being social' to 'being able to serve'.
Figure 6: HuiLiKe's self-developed 'StarReturn' series of biomimetic interactive humanoid robots
From an industry perspective, the value of the L0 to L5 framework provides a set of comparable capability coordinates.
It clearly defines the tasks the system is responsible for, the conditions under which it operates, and the stages where human intervention is required, allowing purchasers, scenario providers, and developers to share a common language. For clients, this is closer to decision-making criteria than 'how human-like it is': whether the robot can complete a closed-loop goal, remember individuals, adapt to scenarios, and take over in case of anomalies. HuiLiKe hopes to continuously deploy, record processes, and conduct professional evaluations to gradually form verifiable and reusable interaction evaluation methods.
In terms of implementation paths, HuiLiKe does not start from general companionship but builds delivery capabilities from professional scenarios. High difficulty scenarios drive technical limits, while deliverable scenarios validate product lower limits. One end aims at L5 non-cooperative strong active interaction as a long-term goal, in high difficulty scenarios such as autism assistance intervention, allowing models to learn from experts on how to find entry points, understand subtle feedback, and adjust strategies; the other end forms products in scenarios with delivery capabilities from L2 to L4, validating multi-round interaction, goal closure, individualized strategies, and scenario coordination capabilities. High difficulty demands define technical directions, while phased applications form commercial delivery, both jointly supporting data and capability iteration.
Therefore, transitioning from interaction frameworks to real service essentially transforms a robot's 'social' capability from a demonstration ability to a service capability: it can both understand humans and adjust strategies, and deliver stably within professional boundaries. Based on the team's long-term technical accumulation, supported by expert methods and real interaction data, HuiLiKe is advancing its thinking on robots' 'social' capabilities into specific technical frameworks and product practices, and continues to conduct verification around multimodal embodied robot interaction capabilities, pushing robots to truly enter thousands of households.
Official contact email: [email protected]
Business cooperation contact person: Mr. Fang (+86) 13308095608 [WeChat same number] [email protected]
END
Industrial Robot Companies
Eston Automation
|
Eve Robot
|
Fao Robot
|
Yuejiang Robot
|
Jieka Robot
|
Songling Robot
|
Luoshi Robot
|
Atongmu Robot
|
Jizhi Jia
|
Hikvision Robot
|
Yifei Technology
|
Ailite Robot
Service and Specialized Robot Companies
Yijiahe
|
Jingpin Tieshang
|
Qiteng Robot
|
Shihé Robot
|
Pudu Robot
|
Shirode Robot
|
Kuma Technology MAMMOTION
Humanoid Robot Companies
优必选科技
|
宇树
|
云深处
|
星动纪元
|
伟景机器人
|
逐际动力
|
乐聚机器人
|
大象机器人
|
魔法原子
|
众擎机器人
|
帕西尼感知
|
赛博格机器人
|
数字华夏
|
傅利叶智能
|
天链机器人
|
开普勒人形机器人
|
灵宝CASBOT
|
清宝机器人
|
浙江人形机器人创新中心
|
动易科技
|
智身科技
|
PNDbotics
|
卓益得机器人
|
擎朗智能
|
伽利略GALILEO
|
松延动力
|
天机智能
|
卧安机器人
|
理工华汇
|
加速进化
具身智能企业
跨维智能
|
银河通用
|
千寻智能
|
灵心巧手
|
睿尔曼智能
|
微亿智造
|
推行科技
|
中科硅纪
|
枢途科技
|
灵巧智能
|
星尘智能
|
穹彻智能
|
Ark Infinite
| iFlytek |
Beijing Humanoid Robot Innovation Center
|
National and Local Co-built Humanoid Robot Innovation Center
|
Daimeng Robot
|
VisiBot Robot
|
StarChart
|
Yuequan Bionics
|
Zeroth Robot
|
Zhongke ShenGu
|
Zhi Fang Square
|
Daka Robot
|
Haocun Technology
|
Ju Shi Intelligence
|
Xynova Xi Nuo Future
|
Feixi Technology
|
Future Power
|
Boden Intelligence
|
Qianjue Technology
|
Lingsheng Technology
|
Ji Cui Intelligent Manufacturing
|
Xinbeit Technology
|
Chen Hunxian Technology
|
Dexmal Original Force Lingji
|
Youliqi
|
Self-variable
|
Ruiyan Intelligent Control Dexterous Hand
|
Qiwu Technology
|
RoboScience Machine Science
|
Zhongke Fifth Dynasty
|
Critical Point
|
Danghong Technology
|
Qiaojie Shu Wu
|
Vbot Weitadongli
|
Tashan Technology
|
Ju Brain磐石
|
Youai Zhizhe Robot
|
Zhi Xing Jian
|
Aimiao Robot
|
Lu Ming Robot
|
Quantum Intelligence
|
Ruoyu Technology
|
Tiehu Robot
|
Zhi Wuji
Medical Robot Enterprises
Yuanhua Intelligence
|
Tianzhihang
|
Sizhe Rui Intelligent Medical
|
Jingfeng Medical
|
Tuo Dao Medical
|
Zhenyida
|
Shu Rui® Robot
|
Luosenbot
|
Shui Mu Dongfang
|
Kangnuositeng
|
Dishimedical
Upstream Industry Chain Enterprises
Lvde Harmonic Wave
|
Yinshi Robot
|
Kunwei Technology
|
Maita Intelligent
|
Qingtong Vision
|
Benmo Technology
|
Lan Dian Touch Control
| Xinjingcheng Sensor
|
BrainCo Qiangbrain Technology
|
Yuli Instrument
|
Ji Ya Jingji
|
Sikan Technology
|
Shenyuan Sheng
|
Feipu Navigation Technology
|
Inks
|
Juxia Intelligent Drive
|
Lingyun Guang Yuanke Vision
|
Xuanji Power
|
Yiyou Technology
| Ruiyuan Precision |
Lingzhu Era
|
HIT Huawei Ke
|
Xinghui Sensor
|
Lingdi Technology
|
Quan Zhi Bo
|
CubeMars Robot Power
|
Wanglong Robot Elevator
|
Hangkai Microelectronics
|
Wutong Sensing Control
Source:RoboSpeak — WeChat · mp.weixin.qq.com