Concept
Vision-language model
A vision-language model (VLM) is a neural network that takes both images and text as input and produces text as output. Think of a large language model, except the context window can hold pixels alongside tokens.
The rest of “Vision-language model” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→