Fig.1

Concept

Mixture-of-Transformers (MoT)

Mixture-of-Transformers (MoT) is a transformer architecture that keeps separate feed-forward and normalization parameters per modality while sharing a single attention path across all of them. Think of it like Mixture-of-Experts, except the routing is fixed by modality…

The rest of “Mixture-of-Transformers (MoT)” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library