VLA

A VLA, or vision-language-action model, is an AI model that takes images and language instructions as input and outputs robot actions.

1 article
Last mentioned

A VLA (vision-language-action) model understands camera images and natural-language instructions together and directly produces robot actions such as joint positions and motion trajectories. A common approach is to train a large vision and language model further on robot action data.

Unlike vision-language models, which process language and images, its output is physical action.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

Gemini Robotics 2 is a VLA, or vision-language-action model: it takes images and language as input and outputs robot actions.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.