Skip to content
Trending eventWatching

VLA-ACL: Action-Consistent Visual Token Pruning for Efficien

1 reports1 reporting sourcesUpdated 2 days ago

Overview

Source roundup

Source roundup from published reports. Claims below are attributed to their publishers, not independently verified. arXiv Robotics — research abstracts: VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models. VLA-ACL is a novel approach that prunes visual tokens in vision-language-action models using action-level supervision, while keeping the base model frozen. It achieves up to 87.5% token pruning, 75% computation reduction…

Generated from attributed reports · Updated 45 minutes ago

Event evidence and corrections

0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.

Reported quantity · inference speedup: 1.5 other · Basis not reported
Supporting report

“s to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These r”

Exact source · revision 1

Source owner not reported

Artifact availability · code: available
Supporting report

“esults establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.”

Exact source · revision 1

Source owner not reported

Report timeline

Follow attributed reports and material updates.

10/9
  1. arXiv Robotics — research abstracts
    VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

    VLA-ACL is a novel approach that prunes visual tokens in vision-language-action models using action-level supervision, while keeping the base model frozen. It achieves up to 87.5% token pruning, 75% computation reduction, and a 1.5x inference speedup on LIBERO and real-world tasks.

Event coverage history

There is not enough continuous observation data to show a trend.

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).