PE Core t16 classify
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
PE Core t16 classify
Image input 384 × 384, text length 32. Classify example: 3 classes, one prompt each. Tensor sizes exclude batch.
LibreYOLO
PE Core t16 classify
Image input 384 × 384, text length 32. Classify example: 3 classes, one prompt each. Tensor sizes exclude batch.
Image tower
Image
3 × 384 × 384
Conv2d 16×16 / 16
192 channels, bias=False
Flatten patch grid
576 × 192
Learned CLS
1 × 192
Concat CLS and patches
577 × 192
Position embedding
577 × 192
+
LayerNorm before encoder
577 × 192
Vision encoder block
577 × 192, repeats=12
LayerNorm after encoder
577 × 192
Latent attention pooling
1 query, 8 pooling heads
Linear projection
192 to 512, bias=True
L2 normalize
512
RoPE applies to patch Q/K in every block.
g14 has no CLS token; other sizes have one.
Text tower
Text token IDs
32 IDs, vocabulary 49,408
Token embedding
32 × 512
Embedded sequence
32 × 512
Position embedding
32 × 512
+
Text transformer block
32 × 512, repeats=12
Final LayerNorm
32 × 512
Select EOT
Argmax token ID, width 512
MatMul projection
512 to 512, no bias
L2 normalize
512
Task output
Image vector
512
Class matrix
3 × 512
MatMul cosine similarity
3 scores
Multiply exp(logit_scale)
3 logits
Prediction applies softmax over the class set.
Image vector: task output
Text vector: task output
Latent attention pooling
Learned latent
1 × 192
Cross-attention
1 query over 577 image tokens
LayerNorm
1 × 192
Linear
192 to 768
GELU
1 × 768
Linear
768 to 192
+
Select pooled token
192
Pooling cross-attention
Query input
1 × 192
Key/value input
577 × 192
Linear Q
192 to 192
Reshape heads
8 × 1 × 24
Linear K
192 to 192
Reshape heads
8 × 577 × 24
Linear V
192 to 192
Reshape heads
8 × 577 × 24
MatMul Q K-transpose
8 × 1 × 577
Scale
Divide by sqrt(24)
Softmax over keys
8 × 1 × 577
MatMul attention × V
8 × 1 × 24
Merge heads
1 × 192
Linear output
192 to 192
Vision encoder block
Input 577 × 192
LayerNorm
577 × 192
Multi-head self-attention
3 heads
+
LayerNorm
577 × 192
Linear
192 to 768
GELU
577 × 768
Linear
768 to 192
+
Output 577 × 192
Vision rotary self-attention
Query input
577 × 192
Key/value input
577 × 192
Linear Q
192 to 192
Reshape heads
3 × 577 × 64
Linear K
192 to 192
Reshape heads
3 × 577 × 64
Linear V
192 to 192
Reshape heads
3 × 577 × 64
Rotary position
1 prefix; rotate patches
Rotary position
1 prefix; rotate patches
MatMul Q K-transpose
3 × 577 × 577
Scale
Divide by sqrt(64)
Softmax over keys
3 × 577 × 577
MatMul attention × V
3 × 577 × 64
Merge heads
577 × 192
Linear output
192 to 192
Text transformer block
Input 32 × 512
LayerNorm
32 × 512
Multi-head self-attention
8 heads
+
LayerNorm
32 × 512
Linear
512 to 2048
GELU
32 × 2048
Linear
2048 to 512
+
Output 32 × 512
Text causal self-attention
Query input
32 × 512
Key/value input
32 × 512
Linear Q
512 to 512
Reshape heads
8 × 32 × 64
Linear K
512 to 512
Reshape heads
8 × 32 × 64
Linear V
512 to 512
Reshape heads
8 × 32 × 64
MatMul Q K-transpose
8 × 32 × 32
Scale
Divide by sqrt(64)
Causal mask
32 × 32
+
Softmax over keys
8 × 32 × 32
MatMul attention × V
8 × 32 × 64
Merge heads
32 × 512
Linear output
512 to 512
Rotary position on Q and K
Q or K per head
576 patch tokens, 64 channels
Prefix bypass
1 CLS token; no rotation
Pair adjacent channels
[a,b] becomes [-b,a]
Cosine table
2D grid offset 1.0
Sine table
Same spatial frequency grid
Multiply input × cos
576 × 64
Multiply rotation × sin
576 × 64
+
Concat preserved CLS and rotated patches
Restore the original token order
No RoPE on V. Frequency base 10,000; XY grid. Family Np = Nv - 1.
Video embedding
Two input frames
2 × 3 × 384 × 384
Image tower per frame
2 × 512, before L2 normalization
Mean frame embeddings
512
L2 normalize once
512
Example uses two frames; encode_video accepts arbitrary frame count F.
No temporal attention is added. Frame embeddings are averaged before normalization.
Variant values
Size
S / P
Nv
Ev / nv / hv
Mv
Et / nt / ht
L
D
t16
384 / 16
577
192 / 12 / 3
768
512 / 12 / 8
32
512
s16
384 / 16
577
384 / 12 / 6
1536
512 / 12 / 8
32
512
b16
224 / 16
197
768 / 12 / 12
3072
1024 / 24 / 16
32
1024
l14
336 / 14
577
1024 / 24 / 16
4096
1024 / 24 / 16
32
1024
g14
448 / 14
1024
1536 / 50 / 16
8960
1280 / 24 / 20
72
1280
S/P: input/patch size. Nv: image tokens. Ev/Et: tower width. nv/nt: block counts. hv/ht: heads. Mv: vision MLP width.
The symbolic graph covers t16, s16, b16, l14. g14 has a separate concrete graph because its CLS path is absent.
Source: libreyolo/models/pe/nn.py. Revision a4d0ecc9e17f.
libreyolo.com