Concept
VLA
VLA stands for Vision-Language-Action model: a policy that takes images (and usually a text instruction) as input and outputs robot actions directly. Think of a vision-language model like a captioner or VQA system, except the output head produces motor commands instead…
The rest of “VLA” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→