Skip to content

Technical Directions

Reinforcement learning Latest coverage

Robotics coverage relating to reinforcement learning.

21 selected reportsLast 30 days 13 recordsAll coverage 104 records

updated

Reinforcement learning Signal

2026-10-09Fri1–20 records
  1. arXiv Robotics — research abstracts85

    SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving

    This paper introduces SDPAD, a fully spike-driven end-to-end planning pipeline for autonomous driving. It converts pre-trained ANN perception stacks into integer-spike form, lifts multi-view images into BEV space using Spike-3D-Lift, and plans through Spike-QFormer, a spiking query transformer. SDPAD achieves high accuracy and low energy consumption, outperforming previous SNN planners and matching mainstream ANN planners in energy efficiency.

  2. arXiv Robotics — research abstracts85

    NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

    NavGPT-3 is a new system that connects language models with action policies for autonomous navigation. It uses a hierarchical runtime architecture to enable reasoning, acting, and monitoring in parallel, with the action policy trained on 19.28M examples. The system achieves state-of-the-art results on R2R-CE and matches human performance on RxR-CE, demonstrating the potential of integrating language models with physical control.

  3. RoboSpeak — WeChat85

    How Can Optimus Generate Revenue Without Any Customers?

    The article examines Tesla's production and financial strategy for Optimus, noting that while manufacturing capacity has improved, the absence of external customers means each robot is treated as a capital expenditure rather than revenue. It contrasts Tesla's internal deployment model with Chinese companies that generate revenue through sales, highlighting the challenges of achieving profitability in the humanoid robotics market.

2026-10-08Thu
  1. RoboSpeak — WeChat83

    Focusing on Autism and Other High-Interaction Scenarios, This Company is Defining Robot 'Social' Interaction

    Wuhan Huilike Robotics Co., Ltd. has developed the VLI multimodal embodied interaction model based on research achievements from a team at Huazhong University of Science and Technology. The company focuses on enhancing robots' social capabilities in real-world scenarios. Through an L0 to L5 interaction capability grading framework, the company defines multiple levels of social capabilities ranging from tool execution to non-cooperative strong active interaction. In specialized scenarios such as autism assistance intervention, Huilike combines expert experience with multimodal data to enable robots to achieve the goals of understanding humans, adjusting strategies, and providing continuous service.

2026-10-07Wed
  1. Robotics — Paper and dataset web discovery83

    Our Journey Building Vision-Language-Action Robotics

    This article details the progression of a robotics team's work from simple imitation learning to developing Vision-Language-Action (VLA) models. It outlines challenges in adapting robotic behavior to new environments, the development of models like SmolVLA and π0.5, and the shift toward simulation-generated data for training. The focus is on building robots that can see, understand, and act autonomously in dynamic settings.

2026-10-05Mon
  1. Machine Heart — Robotics on WeChat85

    Can Embodied Intelligence Be Trusted? Sergey Levine Points the Way

    At the 2026 World Robotics Conference, Sergey Levine pointed out that the robotics industry is undergoing a shift from controlling the body to developing decision-making capabilities. He emphasized that robots need to combine prior knowledge with real-world experience to achieve true intelligence. Although robots are still in the foundational technology stage, through data flywheels and scaling, they may eventually enable broader applications.

2026-10-04Sun
  1. Robotics — Paper and dataset web discovery85

    Diffusion models in robotics: a comprehensive review

    This paper presents a comprehensive review of diffusion models in robotics, focusing on their application in robot learning, data scaling, reinforcement learning, and imitation learning. The authors conducted a systematic literature review, selecting peer-reviewed articles and high-impact preprints to analyze the current state of diffusion model research in robotics. The paper discusses various applications of diffusion models, including semantic-level data augmentation, cross-view and morphology synthesis, automated simulation asset and task generation, trajectory optimization, policy representation, hierarchical planning, and real-time deployment. It also addresses the challenges and limitations of current approaches, such as computational cost, inference latency, and data efficiency, while highlighting promising future directions for research and development in the field.

2026-10-01Thu
  1. RoboSpeak — WeChat86

    Motors Don't Need Upgrades, Robots Can Run Faster! Beihang University Science Subjournal Reveals New Logic for Quadruped Locomotion

    Professor Shi Qing's team from Beijing Institute of Technology was inspired by the high-speed running mechanism of the elephant shrew. They designed a micro quadruped robot named FLEXOR equipped with a dual-joint coupled spine. Through dynamic spine-leg coordination, the robot achieved a 31.6% increase in speed and a 32.2% reduction in energy consumption without upgrading its motors. The study reveals the critical role of the timing coordination between the spine and legs in locomotion performance and verifies the universality of this mechanism across different robot morphologies, providing new insights for highly mobile embodied intelligent robots.

2026-09-30Wed
  1. Tech Xplore — Robotics85

    Roboticist says safety should be top concern with humanoid robots

    Aaron Ames, a leading roboticist at Caltech, argues that safety must be the primary concern for humanoid robots. He emphasizes the need for formal mathematical guarantees, such as control barrier functions (CBFs), to ensure safe operation. Ames highlights the gap between impressive demonstrations and real-world safety, urging the industry to prioritize safety standards to prevent harm.

  2. QbitAI — Robotics85

    GPT-6 Astra Connects to Unitree G1, Cleans the Kitchen!

    The Stanford team's HomeBody project demonstrates how GPT-6 Astra controls the Unitree G1 robot to complete kitchen cleaning tasks through a three-step process: spatial exploration, digital twin simulation, and task execution. The system uses pre-trained models and a skill library to perform complex operations in new environments without additional training.

  3. RoboSpeak — WeChat86

    Humanoid Robot Patents: China Claims 60% Share, What Next?

    The article analyzes China's leading position in the field of humanoid robot patents, noting that as of the first half of 2026, China holds 60.8% of relevant patents, significantly surpassing the United States and South Korea. The article examines the evolution of patent strategies, the trend of technological shift from hardware to software, and the progress made by Chinese companies in patent quality, application scenarios, and global layout. It also points out that patent quantity does not equate to comprehensive technological leadership, emphasizing the challenges of manufacturing costs and quality, as well as the dual role of patents in industrial competition.

2026-09-22Tue
2026-09-14Mon
  1. MIT — Robotics85

    New method enables AI for safety-critical situations

    MIT researchers have developed a new technique that helps generative AI models meet strict safety and task-specific constraints without sacrificing output quality. The method, called HardFlow, reformulates constraint satisfaction as a trajectory-optim, allowing models to explore more freely during generation while ensuring final compliance with hard constraints. It is tested on robotics, control, and computer vision tasks, consistently outperforming existing methods in constraint satisfaction and solution quality.

2026-07-28Tue
2026-07-22Wed
  1. NVIDIA — Robotics85

    NVIDIA Open Sources First GPU-Accelerated Medical Physics Simulation Framework

    NVIDIA has released an open-source, GPU-accelerated Medical Physics Simulation framework as part of its Isaac for Healthcare platform. This tool enables developers to model anatomy-device interactions, generate complex scenarios, and train robot policies in simulation, significantly reducing development time and improving regulatory readiness.

2026-07-15Wed
  1. Boston Dynamics — Blog85

    Deploying Robots Into (and Onto) the Field

    Boston Dynamics deployed their production Atlas robot at the FIFA World Cup 2026™ Norway-Brazil match, where it performed a ceremonial role and showcased its capabilities. The demonstration involved dynamic movement and precision, trained through reinforcement learning, and highlighted the robot's potential for industrial applications beyond entertainment.

2025-11-13Thu
  1. Dexterity — Blog88

    Transactable World Models

    Dexterity introduces 'Transactable World Models' as a foundational component for Physical AI, emphasizing interpretability, physics integration, and real-time consistency. The approach treats world models as operators rather than data stores, enabling multi-agent coordination and robust manipulation in complex environments. This contrasts with traditional world models that focus on data compression and replay, lacking causal reasoning and physical grounding.

2025-10-24Fri
2025-04-01Tue
  1. Sanctuary AI — Research and announcements65

    Sanctuary AI Leads the Industry in Controlling Advanced Hydraulic Hands Using Reinforcement Learning

    Sanctuary AI has demonstrated advanced manipulation skills using reinforcement learning to train dexterous policies for their high-performance hydraulic hands. The system successfully reorients objects in the real world under gravity and with added weight, showcasing the effectiveness of their approach. This represents the first successful transfer of such policies to five-fingered hands, highlighting the company's proprietary technology and hardware capabilities.

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).