Fig.1

Concept

multimodal LLMs

Multimodal LLMs (MLLMs) are large language models that accept inputs beyond text (images, video, audio) and reason over them in a shared representation, then emit text. Think of a standard text LLM, except the token stream can start with pixels or audio frames instead of…

The rest of “multimodal LLMs” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library