Fig.1

Concept

SigLIP-2

SigLIP-2 (Sigmoid Loss for Language-Image Pre-training, version 2) is a vision encoder from Google that produces image embeddings aligned to text, like CLIP, except it swaps CLIP's softmax contrastive loss for a pairwise sigmoid loss. In CLIP, each image-text similarity is…

The rest of “SigLIP-2” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

SigLIP-2, explained · Fig. 1