CLIP b32 classify
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
CLIP b32 classify
Image input 224 × 224; text length 77. Classify example uses 3 classes and one prompt per class.
LibreYOLO
CLIP b32 classify
Image input 224 × 224; text length 77. Classify example uses 3 classes and one prompt per class.
Image tower
Image
3 × 224 × 224
Conv2d 32×32 / 32
768 × 7 × 7, no bias
Flatten and transpose
49 × 768
Learned CLS
1 × 768
Concat CLS and patches
50 × 768
Position embedding
50 × 768
+
LayerNorm before blocks
50 × 768
Vision transformer block
50 × 768, n=12
LayerNorm after blocks
50 × 768
Select CLS
768
MatMul learned projection
768 to 512, no bias
L2 normalize
512
Text tower
Text token IDs
77 IDs, vocabulary 49,408
Token embedding
77 × 512
Embedded sequence
77 × 512
Position embedding
77 × 512
+
Text transformer block
77 × 512, n=12
LayerNorm
77 × 512
Select EOT position
Argmax token ID, width 512
MatMul learned projection
512 to 512, no bias
L2 normalize
512
Task output
Normalized image
512
Normalized class matrix
3 × 512
MatMul cosine similarities
3 class scores
Multiply exp(logit_scale)
3 logits
Prediction postprocess applies softmax across the configured class set.
Normalized image: task output
Normalized text: task output
Vision transformer block
Input 50 × 768
LayerNorm
50 × 768
Multi-head self-attention
12 heads
+
LayerNorm
50 × 768
Linear
768 to 3072
GELU
50 × 3072
Linear
3072 to 768
+
Output 50 × 768
Vision self-attention
Query input
50 × 768
Key/value input
50 × 768
Linear Q
768 to 768
Reshape heads
12 × 50 × 64
Linear K
768 to 768
Reshape heads
12 × 50 × 64
Linear V
768 to 768
Reshape heads
12 × 50 × 64
MatMul Q K-transpose
12 × 50 × 50
Scale
Divide by sqrt(64)
Softmax over keys
12 × 50 × 50
MatMul attention × V
12 × 50 × 64
Merge heads
50 × 768
Linear output
768 to 768
Text transformer block
Input 77 × 512
LayerNorm
77 × 512
Multi-head self-attention
8 heads
+
LayerNorm
77 × 512
Linear
512 to 2048
GELU
77 × 2048
Linear
2048 to 512
+
Output 77 × 512
Text causal self-attention
Query input
77 × 512
Key/value input
77 × 512
Linear Q
512 to 512
Reshape heads
8 × 77 × 64
Linear K
512 to 512
Reshape heads
8 × 77 × 64
Linear V
512 to 512
Reshape heads
8 × 77 × 64
MatMul Q K-transpose
8 × 77 × 77
Scale
Divide by sqrt(64)
Causal mask
77 × 77
+
Softmax over keys
8 × 77 × 77
MatMul attention × V
8 × 77 × 64
Merge heads
77 × 512
Linear output
512 to 512
Variant values
Size
P
Ev
nv
hv
Et
nt
ht
b32
32
768
12
12
512
12
8
b16
16
768
12
12
512
12
8
l14
14
1024
24
16
768
12
12
D: 512 for b32/b16, 768 for l14. Nv: 50, 197, 257 respectively.
Ev/Et: vision/text width. nv/nt: repeats. hv/ht: heads. P: patch size.
Concrete views contain all resolved widths, head counts and token counts.
Source: libreyolo/models/clip/nn.py. Revision a4d0ecc9e17f.
libreyolo.com