DINOv2 n semantic
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
DINOv2 n semantic
Semantic, 518 × 518 input. DINOv2-S: 384 channels, 12 blocks, 6 heads. Tensor sizes exclude batch.
LibreYOLO
DINOv2 n semantic
Semantic, 518 × 518 input. DINOv2-S: 384 channels, 12 blocks, 6 heads. Tensor sizes exclude batch.
DINOv2-S encoder
Input
3 × 518 × 518
Normalize RGB
ImageNet mean/std
Conv2d 14×14 / 14
384 × 37 × 37, bias=True
Flatten patch grid
1369 × 384
Learned CLS
1 × 384
Concat CLS and patches
1370 × 384
Position embedding
1,370 × 384
Bicubic interpolation
37×37 unchanged
+
Encoder blocks 1 to 3
1370 × 384, repeats=3
Tap stage 3
384 × 37 × 37
Encoder blocks 4 to 6
1370 × 384, repeats=3
Tap stage 6
384 × 37 × 37
Encoder blocks 7 to 9
1370 × 384, repeats=3
Tap stage 9
384 × 37 × 37
Encoder blocks 10 to 12
1370 × 384, repeats=3
Tap stage 12
384 × 37 × 37
One attention window; zero register tokens.
Taps apply LayerNorm, remove CLS, reshape.
Encoder block
Input 1370 × 384
LayerNorm
1370 × 384, eps=1e-6
Self-attention
6 heads, 64 channels per head
Multiply layer scale
384 learned channel weights
+
LayerNorm
1370 × 384
Linear
384 to 1,536
GELU
1370 × 1,536
Linear
1,536 to 384
Multiply layer scale
384 learned channel weights
+
Output 1370 × 384
Encoder self-attention
Query input
1370 × 384
Key/value input
1370 × 384
Linear Q
384 to 384
Reshape heads
6 × 1370 × 64
Linear K
384 to 384
Reshape heads
6 × 1370 × 64
Linear V
384 to 384
Reshape heads
6 × 1370 × 64
MatMul Q K-transpose
6 × 1370 × 1370
Scale
Divide by sqrt(64)
Softmax over keys
6 × 1370 × 1370
MatMul attention × V
6 × 1370 × 64
Merge heads
1370 × 384
Linear output
384 to 384
Projector P4
Tap stage 3
384 × 37 × 37
Tap stage 6
384 × 37 × 37
Tap stage 9
384 × 37 × 37
Tap stage 12
384 × 37 × 37
Concat four selected features
1,536 × 37 × 37
Conv2d 1×1
1,536 to 256, bias=False
Channel LayerNorm
256 × 37 × 37
SiLU
256 × 37 × 37
Split channels
a: 128 channels; b: 128 channels
Bottleneck (no shortcut)
128 × 37 × 37
Bottleneck (no shortcut)
128 × 37 × 37
Bottleneck (no shortcut)
128 × 37 × 37
Concat a, b, bottleneck 1, 2, 3
640 × 37 × 37
Conv2d 1×1, channel LayerNorm, SiLU
640 to 256; output 256 × 37 × 37
The projection ConvX is expanded in the definition to the right.
Projector Bottleneck
Conv2d 3×3 / 1
128 to 128, p=1, bias=False
Channel LayerNorm
128 × 37 × 37
SiLU
128 × 37 × 37
Conv2d 3×3 / 1
128 to 128, p=1, bias=False
Channel LayerNorm
128 × 37 × 37
SiLU
128 × 37 × 37
No residual addition in this C2f configuration.
Projector output ConvX
Conv2d 1×1
640 to 256, p=0, bias=False
Channel LayerNorm
256 × 37 × 37
SiLU
256 × 37 × 37
Final channel LayerNorm
256 × 37 × 37
Task head
Projector P4 output
256 × 37 × 37
Lateral Conv2d 1×1
256 to 256, bias=True
Smoothing Conv2d 3×3
256 to 256, p=1, bias=True
GroupNorm
32 groups, 256 channels
GELU
256 × 37 × 37
Dropout2d (eval identity)
p=0.1
Predict Conv2d 1×1
256 to 19 class logits
Bilinear resize
19 × 518 × 518, align_corners=False
C2f repeats three Bottlenecks. Each receives the preceding output; every intermediate feature is concatenated.
Only P4 is configured in all four sizes, so semantic fusion contains one lateral and no multi-level additions.
Size equivalence
Size
Encoder width / depth / heads
Projector width
Selected layers
n
384 / 12 / 6
256
3, 6, 9, 12
s
384 / 12 / 6
256
3, 6, 9, 12
m
384 / 12 / 6
256
3, 6, 9, 12
l
384 / 12 / 6
256
3, 6, 9, 12
All fields used by these heads are identical across sizes at the pinned revision; detector decoder settings do not participate.
Source: libreyolo/models/dinov2/model.py. Revision a4d0ecc9e17f.
libreyolo.com