Fig.1

Concept

Vision-language model

A vision-language model (VLM) is a neural network that takes both images and text as input and produces text as output. Think of a large language model, except the context window can hold pixels alongside tokens.

The rest of “Vision-language model” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

Vision-language model, explained · Fig. 1