EoMT s semantic

Click a block to read its description, or select it with Tab and Enter.

EoMT s semanticSemantic, 512 × 512 input, 100 queries, 150 classes. Query tokens enter only the final 3 encoder blocks.LibreYOLOEoMT s semanticSemantic, 512 × 512 input, 100 queries, 150 classes. Query tokens enter only the final 3 encoder blocks.Encoder-only mask transformerNormalize RGBImageNet mean/stdConv2d 16×16 / 16384 × 32 × 32Flatten patch tokens1024 × 384Add patch position embedding1024 × 384Concat CLS + four register tokens + patches1029 × 384Plain encoder blocks1029 × 384, repeats=9Learned queries100 × 384Prepend query tokens1129 × 384Mask-guided encoder blocks1129 × 384, repeats=3Final LayerNorm1129 × 384Prediction heads100 class rows and 100 mask mapsNo DETR decoder or cross-attention module.The shared encoder attends jointly toqueries, five prefix tokens and image patches.Encoder blockTokens: 1029 before queries, 1129 after.LayerNorm384 channelsSelf-attention6 heads, head width 64Multiply layer scale384 learned channel weights+LayerNorm384 channelsLinear384 to 1536GELU1536 channelsLinear1536 to 384Multiply layer scale384 learned channel weights+Plain encoder self-attentionQuery input1029 × 384Key/value input1029 × 384Linear Q384 to 384Reshape heads6 × 1029 × 64Linear K384 to 384Reshape heads6 × 1029 × 64Linear V384 to 384Reshape heads6 × 1029 × 64MatMul Q K-transpose6 × 1029 × 1029ScaleDivide by sqrt(64)Softmax over keys6 × 1029 × 1029MatMul attention × V6 × 1029 × 64Merge heads1029 × 384Linear output384 to 384Joint masked self-attentionQuery input1129 × 384Key/value input1129 × 384Linear Q384 to 384Reshape heads6 × 1129 × 64Linear K384 to 384Reshape heads6 × 1129 × 64Linear V384 to 384Reshape heads6 × 1129 × 64MatMul Q K-transpose6 × 1129 × 1129ScaleDivide by sqrt(64)Query mask6 × 1129 × 1129+Softmax over keys6 × 1129 × 1129MatMul attention × V6 × 1129 × 64Merge heads1129 × 384Linear output384 to 384Mask guidance before each final blockCurrent joint token state1129 × 384LayerNorm1129 × 384Shared prediction heads100 mask logits at 128 × 128Bilinear resize mask logits100 × 32 × 32Threshold at zero100 × 1024 allowed query-to-patch entriesBuild additive attention mask6 × 1129 × 1129; disallowed=-1e9Other query/key pairs remain allowed. Masking is applied when attn_mask_probs for the block is positive.Fresh configurations use ones. Checkpoint buffer values can disable this guidance; the encoder path stays present.The next block consumes this mask and the current token state, then the prediction is recomputed.Shared prediction headsNormalized joint tokens1129 × 384Select query tokens100 × 384Remove queries, CLS and registers1024 × 384Linear class predictor384 to 151, including no-objectMask embedding MLP100 × 384Reshape and two upscale blocks384 × 128 × 128Einsum query embeddings × pixel embeddings100 × 128 × 128 mask logitsThe same heads serve mask guidance and the final predictions.Mask embedding MLPLinear384 to 384GELU100 × 384Linear384 to 384GELU100 × 384Linear384 to 384Two upscale blocksConvTranspose2d 2×2 / 2384 to 384; 64 × 64GELU384 × 64 × 64Depthwise Conv2d 3×3g=384, p=1, bias=FalseChannel LayerNorm384 × 64 × 64ConvTranspose2d 2×2 / 2384 to 384; 128 × 128GELU384 × 128 × 128Depthwise Conv2d 3×3g=384, p=1, bias=FalseChannel LayerNorm384 × 128 × 128LibreYOLO task outputBilinear resize mask logits100 × 512 × 512, align_corners=FalseClass logits from class predictor100 × 151Softmax, discard no-object100 × 150Sigmoid mask logits100 × 512 × 512Einsum class probabilities × mask probabilities150 × 512 × 512Argmax across semantic classes512 × 512 labelsPostprocessing is outside the learned encoder and heads.Variant valuesSizeEnbhMLP widths38412361536b768123123072l1024244164096Source: libreyolo/models/eomt/nn.py. Revision a4d0ecc9e17f.libreyolo.com