Understanding Persistence in 3D Object Memory from Egocentric Videos
Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
Ledger improves object memory accuracy by tracking locations and histories. It enhances HD-EPIC and UCS-Bench scores and localizes Ego4D objects with high precision.
Source: arXiv Robotics — research abstracts · Read original article ↗
Loading article text…
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
Reported numbers
Reported accuracy
29.7%
Reported accuracy
42.6%
Reported accuracy
33.8%
View original evidence
Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5%
Open source S4Reported accuracy
38.5%
View original evidence
Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5%
Open source S4median error
0.99
View original evidence
localizes Ego4D objects with a 0.99 m median error on returned predictions
Open source S4number of streams
100
View original evidence
Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction
Open source S5
What remains unknown
Not established in the collected evidence: Environment, Control, Data origin.
Reported performance applies to the described task. It does not establish general autonomy or deployment readiness.
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
ion noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementar
Open source S4
y roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.
Open source S5
Source:arXiv Robotics — research abstracts · arxiv.org