a lightweight VLA extended with real-time audio perception
audio-smolvla explores whether a small vision-language-action model can react to sound without losing its pretrained visuomotor abilities. we aligned a pretrained audio encoder with SmolLM2’s embedding space, extended LeRobot to collect and train on audio, and tested the system on a custom SO-101 manipulation task where a bell changes the robot’s next action. reach out if you want to read the paper or talk multimodal robotics.