Fig.1

Concept

Vision Transformer

A Vision Transformer (ViT) applies the standard Transformer encoder, the same architecture used for text, to image classification with almost no vision-specific machinery. If you know how a Transformer processes a sequence of word tokens, ViT is the same thing, except the…

The rest of “Vision Transformer” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

Vision Transformer, explained · Fig. 1