Skip to content
arXiv Robotics — research abstracts· Owen Du, Yang Yue, Jie Zhang, Jiaqi Pi, Chi Bene Chen, Gao Huang·· 2 days agoEditorial score65

VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

Summary

VLA-ACL is a novel approach that prunes visual tokens in vision-language-action models using action-level supervision, while keeping the base model frozen. It achieves up to 87.5% token pruning, 75% computation reduction, and a 1.5x inference speedup on LIBERO and real-world tasks.

Source: arXiv Robotics — research abstracts · Read original article ↗

Loading article text…

What the source reports

Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.

Reported numbers

  • inference speedup

    1.5

    View original evidence
    achieves a 1.5x inference speedup
    Open source S4
Source excerpts and review record

Automatically extracted; no manual editorial approval recorded.

s to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These r

Open source S4

esults establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.

Open source S5

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).