ViT ti

Click a block to read its description, or select it with Tab and Enter.

ViT tiClassification, 224 × 224 input, patch size 16, 1,000 classes. Token shapes exclude batch.LibreYOLOViT tiClassification, 224 × 224 input, patch size 16, 1,000 classes. Token shapes exclude batch.NetworkInput3 × 224 × 224Conv2d 16×16 / 16192 × 14 × 14, bias=TrueFlatten and transpose196 × 192CLS token1 × 192Concat tokens197 × 192Position197 × 192+Transformer block197 × 192, n=12LayerNorm197 × 192, eps=1e-6Select CLS token192Linear classifier192 to 1,000 logitsLearned CLS and absolute position.Dropout probabilities are zero.All Linear projections include bias.Transformer blockInput: 197 × 192LayerNorm197 × 192Self-attention3 heads, width 64 each+LayerNorm197 × 192Linear192 to 768GELU197 × 768Linear768 to 192+Output: 197 × 192Self-attentionLinear QKV192 to 576Reshape and split Q, K, VEach: 3 × 197 × 64Q3 × 197 × 64K3 × 197 × 64V3 × 197 × 64MatMul Q K-transpose3 × 197 × 197ScaleDivide by sqrt(64)Softmax over keys3 × 197 × 197MatMul attention × V3 × 197 × 64Transpose and merge heads197 × 192Linear output projection192 to 192SDPA is drawn as its primitive equation.No causal mask or relative position bias.Variant valuesSizeE: token widthn: block repeatsh: headsMLP widthHead widthti19212376864s384126153664b7681212307264l10242416409664Source: libreyolo/models/vit/nn.py. Revision a4d0ecc9e17f.libreyolo.com