Fig.1

Concept

Parallel Box Decoding

Parallel Box Decoding (PBD) predicts all coordinates of a bounding box in one forward step instead of emitting them one token at a time. Standard vision-language grounding models treat a box like text: an autoregressive decoder produces x1, then y1, then x2, then y2,…

The rest of “Parallel Box Decoding” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library