Concept
multimodal LLMs
Multimodal LLMs (MLLMs) are large language models that accept inputs beyond text (images, video, audio) and reason over them in a shared representation, then emit text. Think of a standard text LLM, except the token stream can start with pixels or audio frames instead of…
The rest of “multimodal LLMs” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→