VLA
A VLA, or vision-language-action model, is an AI model that takes images and language instructions as input and outputs robot actions.
A VLA (vision-language-action) model understands camera images and natural-language instructions together and directly produces robot actions such as joint positions and motion trajectories. A common approach is to train a large vision and language model further on robot action data.
Unlike vision-language models, which process language and images, its output is physical action.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.