Fig.1

Concept

Vision-Language-Action model

A Vision-Language-Action model (VLA) is a policy network that maps images and a natural-language instruction directly to robot actions. Think of a vision-language model like LLaVA, except the output head emits motor commands (joint positions, gripper deltas, end-effector…

The rest of “Vision-Language-Action model” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

Vision-Language-Action model, explained · Fig. 1