Skip to content
Source
arXiv Robotics — research abstracts· Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu·· 2 hours agoEditorial score68

ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

Summary

This paper presents ROMA, an LLM-based system for real-world object-centric multi-sensory active perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop, enabling the model to identify missing evidence and determine interactions. The authors construct ROMI-2K, a large-scale dataset of 2,000 objects and 6 atomic interactions, and develop a two-stage training framework to align sensory modalities and enable reasoning over multi-sensory feedback. Experiments show ROMA can actively acquire missing evidence and solve complex, long-chain perception tasks.

Source: arXiv Robotics — research abstracts · Read original article ↗

Article text · Original source · English

arXiv:2610.06955v1 Announce Type: new Abstract: Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).