SAM 3 large visual prompts (default config)
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
SAM 3 large visual prompts (default config)
Transformers 5.16.1 default architecture, 1008 × 1008 input, one box. Gated checkpoint configuration was not accessible.
LibreYOLO
SAM 3 large visual prompts (default config)
Transformers 5.16.1 default architecture, 1008 × 1008 input, one box. Gated checkpoint configuration was not accessible.
Image-mode visual path
Image encoder
32-block ViT, 1,024 channels
Independent FPN resamplings
256-channel maps at 288,144,72,36
Select three finest feature levels
288×288, 144×144, 72×72
Project high-resolution features
32 × 288 × 288 and 64 × 144 × 144
Add no-memory embedding
256 × 72 × 72 main image feature
No video memory runs in this adapter.
Box prompt encoder
One box
2 corner coordinates × 2
Shift pixel centers
Add 0.5 to x and y
Normalize coordinates
Divide by 1008, map to [-1,1]
MatMul random Fourier matrix
2 coords to 128 frequencies
Multiply by 2 pi
2 × 128
Sin
2 × 128
Cos
2 × 128
Concat sin and cos
2 × 256
Corner type
2 learned × 256
+
Sparse prompt embeddings
2 × 256
Learned no-mask embedding
Broadcast to 256 × 72 × 72
Selected input: one box, no mask prompt.
Points use the same Fourier encoding, plus
positive/negative or not-a-point embeddings.
Dense image position: this encoding evaluated
on the centers of the 72 × 72 feature grid.
Mask decoder inputs
Image embedding
256 × 72 × 72
Dense prompt
256 × 72 × 72
+
Flatten image features
5,184 × 256
Learned output tokens
Object + IoU + 4 mask = 6 × 256
Sparse prompts
2 × 256
Concat output and prompt tokens
8 × 256
Two-way block 1
8 queries, 5,184 image tokens
Two-way block 2
8 queries, 5,184 image tokens
Final token-to-image attention
Q+Q0, K+P, V=updated image
+
Final LayerNorm on queries
8 × 256
Q0 is the original token sequence.
P is fixed Fourier image position.
5,184 image tokens feed block 1 keys/values.
Masks and quality outputs
Updated query tokens
1 IoU + 4 mask tokens + 1 object token
Updated image tokens
5,184 × 256
Reshape to image
256 × 72 × 72
ConvTranspose2d 2×2 / 2
256 to 64; 64 × 144 × 144
Channel LayerNorm
64 × 144 × 144
GELU
64 × 144 × 144
ConvTranspose2d 2×2 / 2
64 to 32; 32 × 288 × 288
GELU
32 × 288 × 288
Select four mask tokens
4 × 256
Four independent Linear layers
Each 256 to 256
ReLU
4 × 256
Four independent Linear layers
Each 256 to 256
ReLU
4 × 256
Four independent Linear layers
Each 256 to 32
MatMul mask coefficients × upscaled image
4 × 32 times 32 × 82,944; reshape to 4 × 288 × 288
Select masks 1 to 3
3 × 288 × 288
Mask logits are resized to the input canvas,
then unpadded/resized to the original image.
multimask_output=True is drawn.
False selects mask/score 0 instead.
Object token: Linear 256, ReLU, Linear 256,
ReLU, Linear 1 produces object-score logit.
+
High-res feature s1
64 × 144 × 144
+
High-res feature s0
32 × 288 × 288
Select IoU token 1
256
Linear
256 to 256
ReLU
256
Linear
256 to 256
ReLU
256
Linear
256 to 4
Sigmoid
4 mask-quality scores
Select scores 1 to 3
3 quality scores
Select object token 0
256
Linear
256 to 256
ReLU
256
Linear
256 to 256
ReLU
256
Linear
256 to 1 object-score logit
Two-way block 1
Query state
8 × 256
Image state
5,184 × 256
Token self-attention
8 heads, 32 channels per head
First self-attention has no PE or residual.
LayerNorm
8 × 256
Original tokens Q0
8 × 256
+
Image position P
5,184 × 256
+
Token-to-image attention
Q=8, K/V=5,184; 8 heads × 16
+
LayerNorm
8 × 256
Linear
256 to 2,048
ReLU
8 × 2,048
Linear
2,048 to 256
+
LayerNorm
8 × 256
Original tokens Q0
8 × 256
+
Image-to-token attention
Q=5,184, K/V=8; 8 heads × 16
+
LayerNorm on image state
5,184 × 256
Updated queries: 8 × 256
Two-way block 2
Query state
8 × 256
Image state
5,184 × 256
Token self-attention
8 heads, 32 channels per head
Original tokens Q0
8 × 256
+
Self-attention Q/K include Q0.
+
LayerNorm
8 × 256
Original tokens Q0
8 × 256
+
Image position P
5,184 × 256
+
Token-to-image attention
Q=8, K/V=5,184; 8 heads × 16
+
LayerNorm
8 × 256
Linear
256 to 2,048
ReLU
8 × 2,048
Linear
2,048 to 256
+
LayerNorm
8 × 256
Original tokens Q0
8 × 256
+
Image-to-token attention
Q=5,184, K/V=8; 8 heads × 16
+
LayerNorm on image state
5,184 × 256
Updated queries: 8 × 256
Token self-attention
Query input
8 × 256
Key/value input
8 × 256
Linear Q
256 to 256
Reshape heads
8 × 8 × 32
Linear K
256 to 256
Reshape heads
8 × 8 × 32
Linear V
256 to 256
Reshape heads
8 × 8 × 32
MatMul Q K-transpose
8 × 8 × 8
Scale
Divide by sqrt(32)
Softmax over keys
8 × 8 × 8
MatMul attention × V
8 × 8 × 32
Merge heads
8 × 256
Linear output
256 to 256
Token-to-image attention
Query input
8 × 256
Key/value input
4096 × 256
Linear Q
256 to 128
Reshape heads
8 × 8 × 16
Linear K
256 to 128
Reshape heads
8 × 4096 × 16
Linear V
256 to 128
Reshape heads
8 × 4096 × 16
MatMul Q K-transpose
8 × 8 × 4096
Scale
Divide by sqrt(16)
Softmax over keys
8 × 8 × 4096
MatMul attention × V
8 × 8 × 16
Merge heads
8 × 128
Linear output
128 to 256
Image-to-token attention
Query input
4096 × 256
Key/value input
8 × 256
Linear Q
256 to 128
Reshape heads
8 × 4096 × 16
Linear K
256 to 128
Reshape heads
8 × 8 × 16
Linear V
256 to 128
Reshape heads
8 × 8 × 16
MatMul Q K-transpose
8 × 4096 × 8
Scale
Divide by sqrt(16)
Softmax over keys
8 × 4096 × 8
MatMul attention × V
8 × 4096 × 16
Merge heads
4096 × 128
Linear output
128 to 256
Image-mode decoder
Q0: original eight output/prompt tokens; P: Fourier image position.
Final cross-attention uses the token-to-image equation, with residual and LayerNorm.
Features s0 and s1 are computed once, then added at the two upscaling steps.
The no-memory embedding is added to the 72×72 image feature.
This is image inference. Video memory attention and memory encoder do not run.
multimask_output=True selects mask/IoU slots 1, 2 and 3.
Single-mask eval can fall back to the best multimask using stability thresholds.
SAM 3 vision encoder
Image
3 × 1,008 × 1,008
Conv2d 14×14 / 14
5,184 × 1,024 patch tokens
Tile learned position embedding
24×24 pretrain grid to 72×72 grid
Add patch position
5,184 × 1,024; no CLS token
ViT blocks, repeats=32
1,024 width, 16 heads, MLP 4,736
Reshape final features
1,024 × 72 × 72
Global blocks (1-based): 8, 16, 24, 32.
All other blocks use nine 24×24 windows.
Head width 64; axial RoPE on Q and K.
Vision ViT block
LayerNorm
72 × 72 × 1,024
Window partition when local
Nine 24×24 windows, no padding
Rotary self-attention
576 local / 5,184 global tokens
Reverse windows
72 × 72 × 1,024
+
LayerNorm
1,024 channels
Linear
1,024 to 4,736
GELU
4,736 channels
Linear
4,736 to 1,024
+
Local vision self-attention
Query input
576 × 1024
Key/value input
576 × 1024
Linear Q
1024 to 1024
Reshape heads
16 × 576 × 64
Linear K
1024 to 1024
Reshape heads
16 × 576 × 64
Linear V
1024 to 1024
Reshape heads
16 × 576 × 64
Rotary position
2D axial; scale 1
Rotary position
2D axial; scale 1
MatMul Q K-transpose
16 × 576 × 576
Scale
Divide by sqrt(64)
Softmax over keys
16 × 576 × 576
MatMul attention × V
16 × 576 × 64
Merge heads
576 × 1024
Linear output
1024 to 1024
Global vision self-attention
Query input
5184 × 1024
Key/value input
5184 × 1024
Linear Q
1024 to 1024
Reshape heads
16 × 5184 × 64
Linear K
1024 to 1024
Reshape heads
16 × 5184 × 64
Linear V
1024 to 1024
Reshape heads
16 × 5184 × 64
Rotary position
2D axial; scale 1/3
Rotary position
2D axial; scale 1/3
MatMul Q K-transpose
16 × 5184 × 5184
Scale
Divide by sqrt(64)
Softmax over keys
16 × 5184 × 5184
MatMul attention × V
16 × 5184 × 64
Merge heads
5184 × 1024
Linear output
1024 to 1024
Vision feature pyramid
Shared final ViT feature
1,024 × 72 × 72
ConvTranspose2d 2×2 / 2
1,024 to 512, 144×144
GELU
512 × 144 × 144
ConvTranspose2d 2×2 / 2
512 to 256, 288×288
Conv2d 1×1
256 to 256
Conv2d 3×3 / 1
256 to 256, p=1
256 × 288 × 288
Shared final ViT feature
1,024 × 72 × 72
ConvTranspose2d 2×2 / 2
1,024 to 512, 144×144
Conv2d 1×1
512 to 256
Conv2d 3×3 / 1
256 to 256, p=1
256 × 144 × 144
Shared final ViT feature
1,024 × 72 × 72
Conv2d 1×1
1024 to 256
Conv2d 3×3 / 1
256 to 256, p=1
256 × 72 × 72
Shared final ViT feature
1,024 × 72 × 72
MaxPool2d 2×2 / 2
1,024 × 36 × 36
Conv2d 1×1
1024 to 256
Conv2d 3×3 / 1
256 to 256, p=1
256 × 36 × 36 (not used by image heads)
These branches are independent resamplings, not top-down additions. Three finest levels enter the image heads.
Default-config scope: architecture verified from Apache-2.0 Transformers classes; no gated checkpoint or license acceptance was used.
Source: libreyolo/models/sam/sam3.py. Revision a4d0ecc9e17f.
libreyolo.com