SmolVLA: A Compact and Efficient Vision-Language-Action Model for Affordable Robotics
SmolVLA is a compact, efficient Vision-Language-Action (VLA) model designed for affordable robotics, trainable on a single GPU and deployable on consumer hardware. It matches the performance of larger VLAs through community-driven data and provides a reference implementation for training and inference.
Editorial context:The release of SmolVLA introduces a compact and efficient Vision-Language-Action model designed for affordable robotics, trainable on a single GPU and deployable on consumer hardware. It matches the performance of larger models through community-driven data and provides a reference implementation for training and inference.