Fig.1

Concept

VLA

VLA stands for Vision-Language-Action model: a policy that takes images (and usually a text instruction) as input and outputs robot actions directly. Think of a vision-language model like a captioner or VQA system, except the output head produces motor commands instead…

The rest of “VLA” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library