DINOv2 n semantic

Click a block to read its description, or select it with Tab and Enter.

DINOv2 n semanticSemantic, 518 × 518 input. DINOv2-S: 384 channels, 12 blocks, 6 heads. Tensor sizes exclude batch.LibreYOLODINOv2 n semanticSemantic, 518 × 518 input. DINOv2-S: 384 channels, 12 blocks, 6 heads. Tensor sizes exclude batch.DINOv2-S encoderInput3 × 518 × 518Normalize RGBImageNet mean/stdConv2d 14×14 / 14384 × 37 × 37, bias=TrueFlatten patch grid1369 × 384Learned CLS1 × 384Concat CLS and patches1370 × 384Position embedding1,370 × 384Bicubic interpolation37×37 unchanged+Encoder blocks 1 to 31370 × 384, repeats=3Tap stage 3384 × 37 × 37Encoder blocks 4 to 61370 × 384, repeats=3Tap stage 6384 × 37 × 37Encoder blocks 7 to 91370 × 384, repeats=3Tap stage 9384 × 37 × 37Encoder blocks 10 to 121370 × 384, repeats=3Tap stage 12384 × 37 × 37One attention window; zero register tokens.Taps apply LayerNorm, remove CLS, reshape.Encoder blockInput 1370 × 384LayerNorm1370 × 384, eps=1e-6Self-attention6 heads, 64 channels per headMultiply layer scale384 learned channel weights+LayerNorm1370 × 384Linear384 to 1,536GELU1370 × 1,536Linear1,536 to 384Multiply layer scale384 learned channel weights+Output 1370 × 384Encoder self-attentionQuery input1370 × 384Key/value input1370 × 384Linear Q384 to 384Reshape heads6 × 1370 × 64Linear K384 to 384Reshape heads6 × 1370 × 64Linear V384 to 384Reshape heads6 × 1370 × 64MatMul Q K-transpose6 × 1370 × 1370ScaleDivide by sqrt(64)Softmax over keys6 × 1370 × 1370MatMul attention × V6 × 1370 × 64Merge heads1370 × 384Linear output384 to 384Projector P4Tap stage 3384 × 37 × 37Tap stage 6384 × 37 × 37Tap stage 9384 × 37 × 37Tap stage 12384 × 37 × 37Concat four selected features1,536 × 37 × 37Conv2d 1×11,536 to 256, bias=FalseChannel LayerNorm256 × 37 × 37SiLU256 × 37 × 37Split channelsa: 128 channels; b: 128 channelsBottleneck (no shortcut)128 × 37 × 37Bottleneck (no shortcut)128 × 37 × 37Bottleneck (no shortcut)128 × 37 × 37Concat a, b, bottleneck 1, 2, 3640 × 37 × 37Conv2d 1×1, channel LayerNorm, SiLU640 to 256; output 256 × 37 × 37The projection ConvX is expanded in the definition to the right.Projector BottleneckConv2d 3×3 / 1128 to 128, p=1, bias=FalseChannel LayerNorm128 × 37 × 37SiLU128 × 37 × 37Conv2d 3×3 / 1128 to 128, p=1, bias=FalseChannel LayerNorm128 × 37 × 37SiLU128 × 37 × 37No residual addition in this C2f configuration.Projector output ConvXConv2d 1×1640 to 256, p=0, bias=FalseChannel LayerNorm256 × 37 × 37SiLU256 × 37 × 37Final channel LayerNorm256 × 37 × 37Task headProjector P4 output256 × 37 × 37Lateral Conv2d 1×1256 to 256, bias=TrueSmoothing Conv2d 3×3256 to 256, p=1, bias=TrueGroupNorm32 groups, 256 channelsGELU256 × 37 × 37Dropout2d (eval identity)p=0.1Predict Conv2d 1×1256 to 19 class logitsBilinear resize19 × 518 × 518, align_corners=FalseC2f repeats three Bottlenecks. Each receives the preceding output; every intermediate feature is concatenated.Only P4 is configured in all four sizes, so semantic fusion contains one lateral and no multi-level additions.Size equivalenceSizeEncoder width / depth / headsProjector widthSelected layersn384 / 12 / 62563, 6, 9, 12s384 / 12 / 62563, 6, 9, 12m384 / 12 / 62563, 6, 9, 12l384 / 12 / 62563, 6, 9, 12All fields used by these heads are identical across sizes at the pinned revision; detector decoder settings do not participate.Source: libreyolo/models/dinov2/model.py. Revision a4d0ecc9e17f.libreyolo.com