Concept
Vision-Language-Action model
A Vision-Language-Action model (VLA) is a policy network that maps images and a natural-language instruction directly to robot actions. Think of a vision-language model like LLaVA, except the output head emits motor commands (joint positions, gripper deltas, end-effector…
The rest of “Vision-Language-Action model” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→