ViT ti
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
ViT ti
Classification, 224 × 224 input, patch size 16, 1,000 classes. Token shapes exclude batch.
LibreYOLO
ViT ti
Classification, 224 × 224 input, patch size 16, 1,000 classes. Token shapes exclude batch.
Network
Input
3 × 224 × 224
Conv2d 16×16 / 16
192 × 14 × 14, bias=True
Flatten and transpose
196 × 192
CLS token
1 × 192
Concat tokens
197 × 192
Position
197 × 192
+
Transformer block
197 × 192, n=12
LayerNorm
197 × 192, eps=1e-6
Select CLS token
192
Linear classifier
192 to 1,000 logits
Learned CLS and absolute position.
Dropout probabilities are zero.
All Linear projections include bias.
Transformer block
Input: 197 × 192
LayerNorm
197 × 192
Self-attention
3 heads, width 64 each
+
LayerNorm
197 × 192
Linear
192 to 768
GELU
197 × 768
Linear
768 to 192
+
Output: 197 × 192
Self-attention
Linear QKV
192 to 576
Reshape and split Q, K, V
Each: 3 × 197 × 64
Q
3 × 197 × 64
K
3 × 197 × 64
V
3 × 197 × 64
MatMul Q K-transpose
3 × 197 × 197
Scale
Divide by sqrt(64)
Softmax over keys
3 × 197 × 197
MatMul attention × V
3 × 197 × 64
Transpose and merge heads
197 × 192
Linear output projection
192 to 192
SDPA is drawn as its primitive equation.
No causal mask or relative position bias.
Variant values
Size
E: token width
n: block repeats
h: heads
MLP width
Head width
ti
192
12
3
768
64
s
384
12
6
1536
64
b
768
12
12
3072
64
l
1024
24
16
4096
64
Source: libreyolo/models/vit/nn.py. Revision a4d0ecc9e17f.
libreyolo.com