EoMT s semantic
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
EoMT s semantic
Semantic, 512 × 512 input, 100 queries, 150 classes. Query tokens enter only the final 3 encoder blocks.
LibreYOLO
EoMT s semantic
Semantic, 512 × 512 input, 100 queries, 150 classes. Query tokens enter only the final 3 encoder blocks.
Encoder-only mask transformer
Normalize RGB
ImageNet mean/std
Conv2d 16×16 / 16
384 × 32 × 32
Flatten patch tokens
1024 × 384
Add patch position embedding
1024 × 384
Concat CLS + four register tokens + patches
1029 × 384
Plain encoder blocks
1029 × 384, repeats=9
Learned queries
100 × 384
Prepend query tokens
1129 × 384
Mask-guided encoder blocks
1129 × 384, repeats=3
Final LayerNorm
1129 × 384
Prediction heads
100 class rows and 100 mask maps
No DETR decoder or cross-attention module.
The shared encoder attends jointly to
queries, five prefix tokens and image patches.
Encoder block
Tokens: 1029 before queries, 1129 after.
LayerNorm
384 channels
Self-attention
6 heads, head width 64
Multiply layer scale
384 learned channel weights
+
LayerNorm
384 channels
Linear
384 to 1536
GELU
1536 channels
Linear
1536 to 384
Multiply layer scale
384 learned channel weights
+
Plain encoder self-attention
Query input
1029 × 384
Key/value input
1029 × 384
Linear Q
384 to 384
Reshape heads
6 × 1029 × 64
Linear K
384 to 384
Reshape heads
6 × 1029 × 64
Linear V
384 to 384
Reshape heads
6 × 1029 × 64
MatMul Q K-transpose
6 × 1029 × 1029
Scale
Divide by sqrt(64)
Softmax over keys
6 × 1029 × 1029
MatMul attention × V
6 × 1029 × 64
Merge heads
1029 × 384
Linear output
384 to 384
Joint masked self-attention
Query input
1129 × 384
Key/value input
1129 × 384
Linear Q
384 to 384
Reshape heads
6 × 1129 × 64
Linear K
384 to 384
Reshape heads
6 × 1129 × 64
Linear V
384 to 384
Reshape heads
6 × 1129 × 64
MatMul Q K-transpose
6 × 1129 × 1129
Scale
Divide by sqrt(64)
Query mask
6 × 1129 × 1129
+
Softmax over keys
6 × 1129 × 1129
MatMul attention × V
6 × 1129 × 64
Merge heads
1129 × 384
Linear output
384 to 384
Mask guidance before each final block
Current joint token state
1129 × 384
LayerNorm
1129 × 384
Shared prediction heads
100 mask logits at 128 × 128
Bilinear resize mask logits
100 × 32 × 32
Threshold at zero
100 × 1024 allowed query-to-patch entries
Build additive attention mask
6 × 1129 × 1129; disallowed=-1e9
Other query/key pairs remain allowed. Masking is applied when attn_mask_probs for the block is positive.
Fresh configurations use ones. Checkpoint buffer values can disable this guidance; the encoder path stays present.
The next block consumes this mask and the current token state, then the prediction is recomputed.
Shared prediction heads
Normalized joint tokens
1129 × 384
Select query tokens
100 × 384
Remove queries, CLS and registers
1024 × 384
Linear class predictor
384 to 151, including no-object
Mask embedding MLP
100 × 384
Reshape and two upscale blocks
384 × 128 × 128
Einsum query embeddings × pixel embeddings
100 × 128 × 128 mask logits
The same heads serve mask guidance and the final predictions.
Mask embedding MLP
Linear
384 to 384
GELU
100 × 384
Linear
384 to 384
GELU
100 × 384
Linear
384 to 384
Two upscale blocks
ConvTranspose2d 2×2 / 2
384 to 384; 64 × 64
GELU
384 × 64 × 64
Depthwise Conv2d 3×3
g=384, p=1, bias=False
Channel LayerNorm
384 × 64 × 64
ConvTranspose2d 2×2 / 2
384 to 384; 128 × 128
GELU
384 × 128 × 128
Depthwise Conv2d 3×3
g=384, p=1, bias=False
Channel LayerNorm
384 × 128 × 128
LibreYOLO task output
Bilinear resize mask logits
100 × 512 × 512, align_corners=False
Class logits from class predictor
100 × 151
Softmax, discard no-object
100 × 150
Sigmoid mask logits
100 × 512 × 512
Einsum class probabilities × mask probabilities
150 × 512 × 512
Argmax across semantic classes
512 × 512 labels
Postprocessing is outside the learned encoder and heads.
Variant values
Size
E
n
b
h
MLP width
s
384
12
3
6
1536
b
768
12
3
12
3072
l
1024
24
4
16
4096
Source: libreyolo/models/eomt/nn.py. Revision a4d0ecc9e17f.
libreyolo.com