ROMA: LLM System for Real-World Object-Centric Multi-Sensory
Overview
ROMA is an LLM-based system integrating vision, audio, tactile, and force sensing for real-world object-centric perception. It enables reasoning-interaction-feedback loops for active perception.
Generated from attributed reports · 57 minutes agoUpdated
Event evidence and corrections
0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.
No current evidence-backed claims. Missing information remains not reported.
Report timeline
Follow attributed reports and material updates.
- arXiv Robotics — research abstractsROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception
This paper presents ROMA, an LLM-based system for real-world object-centric multi-sensory active perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop, enabling the model to identify missing evidence and determine interactions. The authors construct ROMI-2K, a large-scale dataset of 2,000 objects and 6 atomic interactions, and develop a two-stage training framework to align sensory modalities and enable reasoning over multi-sensory feedback. Experiments show ROMA can actively acquire missing evidence and solve complex, long-chain perception tasks.
Event attention history
There is not enough continuous observation data to show a trend.
Timezone · UTC
Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).