Depth Anything 3 Large
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
Depth Anything 3 Large
DA3MONO-LARGE, inverse-depth output, input 3 × 504 × 504, native eval. Shapes exclude batch.
LibreYOLO
Depth Anything 3 Large
DA3MONO-LARGE, inverse-depth output, input 3 × 504 × 504, native eval. Shapes exclude batch.
Single-view DINOv2-Large
ImageNet normalization + singleton view
B × 1 × 3 × 504 × 504
Conv2d patch14, stride14
1024 × 36 × 36;1296 patch tokens
CLS + interpolated learned positions
1297 × 1024; position grid37×37 resized36×36
Transformer blocks, n=5
Ends at zero-based index4;16 heads
LayerNorm + remove CLS
T1: 1024 × 36 × 36
Transformer blocks, n=7
Ends at zero-based index11;16 heads
LayerNorm + remove CLS
T2: 1024 × 36 × 36
Transformer blocks, n=6
Ends at zero-based index17;16 heads
LayerNorm + remove CLS
T3: 1024 × 36 × 36
Transformer blocks, n=6
Ends at zero-based index23;16 heads
LayerNorm + remove CLS
T4: 1024 × 36 × 36
alt_start=qknorm_start=rope_start=-1.
No alternating view attention, QK norm, RoPE or camera token.
cat_token=False; tap width stays1024.
DPT adapters and fusion
Tap T1
1024 × 36 × 36
Conv2d 1×1
1024 to 256
ConvTranspose2d 4×4
s4,p0; 256 channels
Conv2d 3×3
256 to 256; 144 × 144; bias=False
L1
256 ch
Tap T2
1024 × 36 × 36
Conv2d 1×1
1024 to 512
ConvTranspose2d 2×2
s2,p0; 512 channels
Conv2d 3×3
512 to 256; 72 × 72; bias=False
L2
256 ch
Tap T3
1024 × 36 × 36
Conv2d 1×1
1024 to 1024
Identity
1024 channels
Conv2d 3×3
1024 to 256; 36 × 36; bias=False
L3
256 ch
Tap T4
1024 × 36 × 36
Conv2d 1×1
1024 to 1024
Conv2d 3×3
s2,p1; 1024 channels
Conv2d 3×3
1024 to 256; 18 × 18; bias=False
L4
256 ch
FeatureFusionBlock 4
256 channels; output 36 × 36
L4
256 channels
FeatureFusionBlock 3
256 channels; output 72 × 72
L3
256 channels
FeatureFusionBlock 2
256 channels; output 144 × 144
L2
256 channels
FeatureFusionBlock 1
256 channels; output 288 × 288
L1
256 channels
L1...L4 are named continuations of the independent adapted feature maps.
First fusion resizes to the exact next grid, including odd-grid rounding.
Transformer block
Input tokens
1297 × 1024
LayerNorm
1024 channels; epsilon1e-5
Multihead self-attention
16 heads; head width64
LayerScale
1024 learned channel scalars
+
LayerNorm
1024 channels; epsilon1e-5
Linear
1024 to 4096
GELU
Linear
4096 to 1024
LayerScale
1024 learned channel scalars
+
Output tokens
1297 × 1024
Self-attention primitives
Fused QKV linear
1024 to 3072; biases enabled
Split Q, K and V
16 heads; 1297 tokens;64 channels/head
Q × transpose(K) / 8
1297 × 1297 per head
Values V
64/head
Softmax over keys
Attention weights × V
16 heads;64 output channels/head
Concat heads
1297 × 1024
Output linear
1024 to 1024
No causal mask. Dropout and stochastic depth are inactive in eval.
Shared output feature and two heads
Conv2d 3×3
256 to128;288 × 288
Bilinear resize
128 × 504 × 504; align_corners=True
Conv2d 3×3
128 to32; s1,p1
ReLU
Conv2d 1×1
32 to1
Exponential
1 × 504 × 504
Conv2d 3×3
128 to32; s1,p1
ReLU
Conv2d 1×1
32 to1
ReLU
1 × 504 × 504
Depth output is positive relative depth.
Sky is ReLU output, not a probability sigmoid.
No confidence head when output_dim=1.
Sky handling and inversion execute after this shared core.
ResidualConvUnit
Input
256 channels
ReLU
Conv2d 3×3
256 to 256; s1,p1,bias=True
ReLU
Conv2d 3×3
256 to 256; s1,p1,bias=True
+
No BatchNorm in these DPT fusion units.
FeatureFusionBlock with lateral input
Top-down feature
256 channels
Lateral feature
256 channels
ResidualConvUnit
256 channels
+
ResidualConvUnit
256 channels
Bilinear resize
Target size shown at each occurrence; align_corners=True
Conv2d 1×1
256 to 256; bias=True
Deepest block has no lateral input and skips the first RCU/add.
Native depth finishing
Non-sky mask
sky <0.3, per image
Count both mask regions
Apply only if each has more than10 pixels
Non-sky depth values
Sample100000 if there are more
99th percentile
Far-depth value from non-sky depth
Replace sky pixels
Use far depth where sky>=0.3
Clamp depth minimum1e-6
Positive denominator
Reciprocal
Inverse relative depth
Remove singleton view
B × 1 × 504 × 504
If counts fail, keep original depth before inversion.
Raw adapters:256×144²,512×72²,1024×36²,1024×18². Fusion outputs36²,72²,144²,288² before the final resize.
Sky statistics are computed independently per batch image; the wrapper never mixes images into one scene.
Source: libreyolo/models/depth_anything3/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com