Fig.1

Concept

MLLM

MLLM stands for Multimodal Large Language Model: an LLM that accepts more than just text as input, most often images alongside a text prompt. Think of a standard decoder-only LLM, except the token stream can include patches from an image encoder projected into the same…

The rest of “MLLM” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

MLLM, explained · Fig. 1