Depth Anything V2 S
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
Depth Anything V2 S
Relative inverse depth, native eval, input 3 × 518 × 518. Shapes exclude batch.
LibreYOLO
Depth Anything V2 S
Relative inverse depth, native eval, input 3 × 518 × 518. Shapes exclude batch.
DINOv2 image encoder
RGB + ImageNet normalization
3 × 518 × 518
Conv2d patch embedding 14×14
3 to 384; stride14; 37 × 37 grid
Prepend CLS + add learned positions
1370 × 384; no register tokens
Transformer blocks, n=3
End at zero-based block 2; 6 heads
LayerNorm; remove CLS
T1: 384 × 37 × 37
Transformer blocks, n=3
End at zero-based block 5; 6 heads
LayerNorm; remove CLS
T2: 384 × 37 × 37
Transformer blocks, n=3
End at zero-based block 8; 6 heads
LayerNorm; remove CLS
T3: 384 × 37 × 37
Transformer blocks, n=3
End at zero-based block 11; 6 heads
LayerNorm; remove CLS
T4: 384 × 37 × 37
Tap normalization does not replace the ongoing token stream.
CLS is returned by the backbone but ignored by this DPT head.
s/b/l use GELU MLP; g uses SwiGLU width4096.
Learned position grid is already37×37 at this input size.
All LayerScale parameters initialize to1.0.
DPT adapters and fusion
Tap T1
384 × 37 × 37
Conv2d 1×1
384 to 48
ConvTranspose2d 4×4
s4,p0; 48 channels
Conv2d 3×3
48 to 64; 148 × 148; bias=False
L1
64 ch
Tap T2
384 × 37 × 37
Conv2d 1×1
384 to 96
ConvTranspose2d 2×2
s2,p0; 96 channels
Conv2d 3×3
96 to 64; 74 × 74; bias=False
L2
64 ch
Tap T3
384 × 37 × 37
Conv2d 1×1
384 to 192
Identity
192 channels
Conv2d 3×3
192 to 64; 37 × 37; bias=False
L3
64 ch
Tap T4
384 × 37 × 37
Conv2d 1×1
384 to 384
Conv2d 3×3
s2,p1; 384 channels
Conv2d 3×3
384 to 64; 19 × 19; bias=False
L4
64 ch
FeatureFusionBlock 4
64 channels; output 37 × 37
L4
64 channels
FeatureFusionBlock 3
64 channels; output 74 × 74
L3
64 channels
FeatureFusionBlock 2
64 channels; output 148 × 148
L2
64 channels
FeatureFusionBlock 1
64 channels; output 296 × 296
L1
64 channels
L1...L4 are named continuations of the independent adapted feature maps.
First fusion resizes to the exact next grid, including odd-grid rounding.
Transformer block
Input tokens
1370 × 384
LayerNorm
384 channels; epsilon1e-6
Multihead self-attention
6 heads; head width64
LayerScale
384 learned channel scalars
+
LayerNorm
384 channels; epsilon1e-6
Linear
384 to 1536
GELU
Linear
1536 to 384
LayerScale
384 learned channel scalars
+
Output tokens
1370 × 384
Self-attention primitives
Fused QKV linear
384 to 1152; biases enabled
Split Q, K and V
6 heads; 1370 tokens;64 channels/head
Q × transpose(K) / 8
1370 × 1370 per head
Values V
64/head
Softmax over keys
Attention weights × V
6 heads;64 output channels/head
Concat heads
1370 × 384
Output linear
384 to 384
No causal mask. Dropout and stochastic depth are inactive in eval.
DPT output head
Conv2d 3×3
64 to 32; 296 × 296
Bilinear resize
518 × 518; align_corners=True
Conv2d 3×3
32 to32; s1,p1
ReLU
Conv2d 1×1
32 to1
ReLU + final ReLU
Nonnegative inverse depth
Relative inverse depth
1 × 518 × 518
ResidualConvUnit
Input
64 channels
ReLU
Conv2d 3×3
64 to 64; s1,p1,bias=True
ReLU
Conv2d 3×3
64 to 64; s1,p1,bias=True
+
No BatchNorm in these DPT fusion units.
FeatureFusionBlock with lateral input
Top-down feature
64 channels
Lateral feature
64 channels
ResidualConvUnit
64 channels
+
ResidualConvUnit
64 channels
Bilinear resize
Target size shown at each occurrence; align_corners=True
Conv2d 1×1
64 to 64; bias=True
Deepest block has no lateral input and skips the first RCU/add.
Source: libreyolo/models/depth_anything/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com