MobileSAM tiny
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
MobileSAM tiny
Promptable segmentation, 1024 × 1024 input, one box, multimask output. TinyViT encoder and native SAM decoder.
LibreYOLO
MobileSAM tiny
Promptable segmentation, 1024 × 1024 input, one box, multimask output. TinyViT encoder and native SAM decoder.
TinyViT image encoder
RGB input after normalize/pad
3 × 1,024 × 1,024
Conv2d 3×3 / 2
3 to 32, 512 × 512
BatchNorm2d
32 × 512 × 512
GELU
32 × 512 × 512
Conv2d 3×3 / 2
32 to 64, 256 × 256
BatchNorm2d
64 × 256 × 256
MBConv, repeats=2
64 × 256 × 256
Patch merge 1
128 × 128 × 128
TinyViT block, repeats=2
16,384 tokens × 128
Patch merge 2
160 × 64 × 64
TinyViT block, repeats=6
4,096 tokens × 160
Patch merge 3 (stride 1)
320 × 64 × 64
TinyViT block, repeats=2
4,096 tokens × 320
Conv2d 1×1 + channel LN
320 to 256, 64 × 64
Conv2d 3×3 + channel LN
256 to 256, p=1, 64 × 64
Box prompt encoder
One box
2 corner coordinates × 2
Shift pixel centers
Add 0.5 to x and y
Normalize coordinates
Divide by 1024, map to [-1,1]
MatMul random Fourier matrix
2 coords to 128 frequencies
Multiply by 2 pi
2 × 128
Sin
2 × 128
Cos
2 × 128
Concat sin and cos
2 × 256
Corner type
2 learned × 256
+
Sparse prompt embeddings
2 × 256
Learned no-mask embedding
Broadcast to 256 × 64 × 64
Selected input: one box, no mask prompt.
Points use the same Fourier encoding, plus
positive/negative or not-a-point embeddings.
Dense image position: this encoding evaluated
on the centers of the 64 × 64 feature grid.
Mask decoder inputs
Image embedding
256 × 64 × 64
Dense prompt
256 × 64 × 64
+
Flatten image features
4,096 × 256
Learned output tokens
IoU + 4 masks = 5 × 256
Sparse prompts
2 × 256
Concat output and prompt tokens
7 × 256
Two-way block 1
7 queries, 4,096 image tokens
Two-way block 2
7 queries, 4,096 image tokens
Final token-to-image attention
Q+Q0, K+P, V=updated image
+
Final LayerNorm on queries
7 × 256
Q0 is the original token sequence.
P is fixed Fourier image position.
4,096 image tokens feed block 1 keys/values.
Masks and quality outputs
Updated query tokens
1 IoU + 4 mask tokens
Updated image tokens
4,096 × 256
Reshape to image
256 × 64 × 64
ConvTranspose2d 2×2 / 2
256 to 64; 64 × 128 × 128
Channel LayerNorm
64 × 128 × 128
GELU
64 × 128 × 128
ConvTranspose2d 2×2 / 2
64 to 32; 32 × 256 × 256
GELU
32 × 256 × 256
Select four mask tokens
4 × 256
Four independent Linear layers
Each 256 to 256
ReLU
4 × 256
Four independent Linear layers
Each 256 to 256
ReLU
4 × 256
Four independent Linear layers
Each 256 to 32
MatMul mask coefficients × upscaled image
4 × 32 times 32 × 65,536; reshape to 4 × 256 × 256
Select masks 1 to 3
3 × 256 × 256
IoU token Linear
256 to 256
ReLU
256
Linear
256 to 256
ReLU
256
Linear
256 to 4 scores
Select scores 1 to 3
3 quality scores
Mask logits are resized to the input canvas,
then unpadded/resized to the original image.
multimask_output=True is drawn.
False selects mask/score 0 instead.
Two-way block 1
Query state
7 × 256
Image state
4,096 × 256
Token self-attention
8 heads, 32 channels per head
First self-attention has no PE or residual.
LayerNorm
7 × 256
Original tokens Q0
7 × 256
+
Image position P
4,096 × 256
+
Token-to-image attention
Q=7, K/V=4,096; 8 heads × 16
+
LayerNorm
7 × 256
Linear
256 to 2,048
ReLU
7 × 2,048
Linear
2,048 to 256
+
LayerNorm
7 × 256
Original tokens Q0
7 × 256
+
Image-to-token attention
Q=4,096, K/V=7; 8 heads × 16
+
LayerNorm on image state
4,096 × 256
Updated queries: 7 × 256
Two-way block 2
Query state
7 × 256
Image state
4,096 × 256
Token self-attention
8 heads, 32 channels per head
Original tokens Q0
7 × 256
+
Self-attention Q/K include Q0.
+
LayerNorm
7 × 256
Original tokens Q0
7 × 256
+
Image position P
4,096 × 256
+
Token-to-image attention
Q=7, K/V=4,096; 8 heads × 16
+
LayerNorm
7 × 256
Linear
256 to 2,048
ReLU
7 × 2,048
Linear
2,048 to 256
+
LayerNorm
7 × 256
Original tokens Q0
7 × 256
+
Image-to-token attention
Q=4,096, K/V=7; 8 heads × 16
+
LayerNorm on image state
4,096 × 256
Updated queries: 7 × 256
Token self-attention
Query input
7 × 256
Key/value input
7 × 256
Linear Q
256 to 256
Reshape heads
8 × 7 × 32
Linear K
256 to 256
Reshape heads
8 × 7 × 32
Linear V
256 to 256
Reshape heads
8 × 7 × 32
MatMul Q K-transpose
8 × 7 × 7
Scale
Divide by sqrt(32)
Softmax over keys
8 × 7 × 7
MatMul attention × V
8 × 7 × 32
Merge heads
7 × 256
Linear output
256 to 256
Token-to-image attention
Query input
7 × 256
Key/value input
4096 × 256
Linear Q
256 to 128
Reshape heads
8 × 7 × 16
Linear K
256 to 128
Reshape heads
8 × 4096 × 16
Linear V
256 to 128
Reshape heads
8 × 4096 × 16
MatMul Q K-transpose
8 × 7 × 4096
Scale
Divide by sqrt(16)
Softmax over keys
8 × 7 × 4096
MatMul attention × V
8 × 7 × 16
Merge heads
7 × 128
Linear output
128 to 256
Image-to-token attention
Query input
4096 × 256
Key/value input
7 × 256
Linear Q
256 to 128
Reshape heads
8 × 4096 × 16
Linear K
256 to 128
Reshape heads
8 × 7 × 16
Linear V
256 to 128
Reshape heads
8 × 7 × 16
MatMul Q K-transpose
8 × 4096 × 7
Scale
Divide by sqrt(16)
Softmax over keys
8 × 4096 × 7
MatMul attention × V
8 × 4096 × 16
Merge heads
4096 × 128
Linear output
128 to 256
Decoder conventions
Q/K add positional terms; V is the raw state.
Final attention repeats token-to-image attention.
Its output is added to queries, then normalized.
Four mask tokens have independent MLP weights.
IoU token uses a separate three-layer MLP.
All decoder Linear projections include bias.
Prompt input for this view is one box.
Dense prompt is the learned no-mask embedding.
Point prompts use Fourier position with a
positive, negative or padding type embedding.
Encode once and cache the image features;
new prompts rerun prompt encoder and decoder.
MBConv
Input: 64 × 256 × 256
Conv2d 1×1
64 to 256, no bias
BatchNorm2d
256 × 256 × 256
GELU
256 × 256 × 256
Depthwise Conv2d 3×3
256 to 256, g=256, p=1
BatchNorm2d
256 × 256 × 256
GELU
256 × 256 × 256
Conv2d 1×1
256 to 64, no bias
BatchNorm2d
64 × 256 × 256
+
GELU
64 × 256 × 256
TinyViT block
Stage order in numeric lists: 128, 160, 320.
Pad and partition windows
128 pads to 133; 64 pads to 70
LayerNorm
128 / 160 / 320 channels
Window self-attention
4 / 5 / 10 heads, 32 channels per head
Reverse windows and crop
128×128 / 64×64 / 64×64
+
Reshape NCHW
128 / 160 / 320 channels
Depthwise Conv2d 3×3
groups=128 / 160 / 320, p=1
BatchNorm2d
128 / 160 / 320 channels
Flatten spatial tokens
16,384 / 4,096 / 4,096 tokens
LayerNorm
128 / 160 / 320 channels
Linear
128 to 512 / 160 to 640 / 320 to 1,280
GELU
512 / 640 / 1,280 channels
Linear
512 to 128 / 640 to 160 / 1,280 to 320
+
Patch merging
Reshape input to NCHW
64×256×256 / 128×128×128 / 160×64×64
Conv2d 1×1
64 to 128 / 128 to 160 / 160 to 320
BatchNorm2d
128 / 160 / 320 channels
GELU
128 / 160 / 320 channels
Depthwise Conv2d 3×3
Strides 2 / 2 / 1, p=1
BatchNorm2d
128 / 160 / 320 channels
GELU
128 / 160 / 320 channels
Conv2d 1×1
128 to 128 / 160 to 160 / 320 to 320
BatchNorm2d
128 / 160 / 320 channels
Flatten and transpose
16,384×128 / 4,096×160 / 4,096×320
No residual connection around patch merging.
Encoder neck primitives:
Conv2d 1×1
320 to 256, bias=False
Channel LayerNorm
256 × 64 × 64
Conv2d 3×3 / 1
256 to 256, p=1, bias=False
Channel LayerNorm
256 × 64 × 64
Window self-attention
Query input
49 / 196 / 49 × 128 / 160 / 320
Key/value input
49 / 196 / 49 × 128 / 160 / 320
Linear Q
128 / 160 / 320 to 128 / 160 / 320
Reshape heads
4 / 5 / 10 × 49 / 196 / 49 × 32
Linear K
128 / 160 / 320 to 128 / 160 / 320
Reshape heads
4 / 5 / 10 × 49 / 196 / 49 × 32
Linear V
128 / 160 / 320 to 128 / 160 / 320
Reshape heads
4 / 5 / 10 × 49 / 196 / 49 × 32
MatMul Q K-transpose
4 / 5 / 10 × 49 / 196 / 49 × 49 / 196 / 49
Scale
Divide by sqrt(32)
Position bias
4 / 5 / 10 × 49 / 196 / 49 × 49 / 196 / 49
+
Softmax over keys
4 / 5 / 10 × 49 / 196 / 49 × 49 / 196 / 49
MatMul attention × V
4 / 5 / 10 × 49 / 196 / 49 × 32
Merge heads
49 / 196 / 49 × 128 / 160 / 320
Linear output
128 / 160 / 320 to 128 / 160 / 320
Attention bias uses absolute relative offsets.
Stage 1: 361 windows of 7×7.
Stage 2: 25 windows of 14×14.
Stage 3: 100 windows of 7×7.
All numeric lists are aligned by the three transformer stages. Their actual dimensions are written in the stage graph and definitions.
Source: libreyolo/models/mobilesam/model.py. Revision a4d0ecc9e17f.
libreyolo.com