Fig.1

Concept

vision-language-action (VLA)

A vision-language-action (VLA) model is a single network that takes images (or video) plus a text instruction and outputs actions, typically robot control commands. Think of it like a vision-language model (VLM) such as LLaVA, except the output head no longer emits text…

The rest of “vision-language-action (VLA)” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library