Concept
VLM
VLM stands for Vision-Language Model: a neural network that takes both images (or video) and text as input and produces text as output. Think of it like an LLM, except the token stream that feeds the transformer can also carry visual patches.
The rest of “VLM” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→