Concept
Vision-language-action models
Vision-language-action models (VLAs) are robot policies that map camera images plus a natural-language instruction directly to robot actions, all inside a single neural network. Think of a vision-language model like the ones behind image chat, except the output head emits…
The rest of “Vision-language-action models” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→