PicoSAM3 pico
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
PicoSAM3 pico
ROI mask segmentation, 96 × 96 input. One mask logit map. Tensor sizes exclude batch.
LibreYOLO
PicoSAM3 pico
ROI mask segmentation, 96 × 96 input. One mask logit map. Tensor sizes exclude batch.
Encoder and decoder
ROI image
3 × 96 × 96
Encoder stage 1
48 × 96 × 96
Conv2d 3×3 / 2
48 × 48 × 48, p=1
Skip Conv2d 1×1
48 to 40, bias=True
Upsample block 4
40 × 96 × 96
+
Encoder stage 2
96 × 48 × 48
Conv2d 3×3 / 2
96 × 24 × 24, p=1
Skip Conv2d 1×1
96 to 80, bias=True
Upsample block 3
80 × 48 × 48
+
Encoder stage 3
160 × 24 × 24
Conv2d 3×3 / 2
160 × 12 × 12, p=1
Skip Conv2d 1×1
160 to 128, bias=True
Upsample block 2
128 × 24 × 24
+
Encoder stage 4
256 × 12 × 12
Conv2d 3×3 / 2
256 × 6 × 6, p=1
Skip Conv2d 1×1
256 to 192, bias=True
Upsample block 1
192 × 12 × 12
+
Bottleneck
320 × 6 × 6
Channel attention (ECA)
40 × 96 × 96
Refine mask logits
1 × 96 × 96
Sigmoid and ROI resize
Mask pasted into original image
Encoder 1
Conv2d 3×3 / 1
3 × 96 × 96; p=1, g=3
BatchNorm2d
3 × 96 × 96
ReLU
3 × 96 × 96
Conv2d 1×1 / 1
48 × 96 × 96; p=0, g=1
BatchNorm2d
48 × 96 × 96
ReLU
48 × 96 × 96
Encoder 2
Conv2d 3×3 / 1
48 × 48 × 48; p=1, g=48
BatchNorm2d
48 × 48 × 48
ReLU
48 × 48 × 48
Conv2d 1×1 / 1
96 × 48 × 48; p=0, g=1
BatchNorm2d
96 × 48 × 48
ReLU
96 × 48 × 48
Encoder 3
Conv2d 3×3 / 1
96 × 24 × 24; p=1, g=96
BatchNorm2d
96 × 24 × 24
ReLU
96 × 24 × 24
Conv2d 1×1 / 1
160 × 24 × 24; p=0, g=1
BatchNorm2d
160 × 24 × 24
ReLU
160 × 24 × 24
Encoder 4
Conv2d 3×3 / 1
160 × 12 × 12; p=1, g=160
BatchNorm2d
160 × 12 × 12
ReLU
160 × 12 × 12
Conv2d 1×1 / 1
256 × 12 × 12; p=0, g=1
BatchNorm2d
256 × 12 × 12
ReLU
256 × 12 × 12
Bottleneck
Conv2d 3×3 / 1
256 × 6 × 6; p=1, g=256
BatchNorm2d
256 × 6 × 6
ReLU
256 × 6 × 6
Conv2d 1×1 / 1
320 × 6 × 6; p=0, g=1
BatchNorm2d
320 × 6 × 6
ReLU
320 × 6 × 6
Conv2d 3×3 / 1
320 × 6 × 6; p=2, g=320, d=2
BatchNorm2d
320 × 6 × 6
ReLU
320 × 6 × 6
Conv2d 1×1 / 1
320 × 6 × 6; p=0, g=1
BatchNorm2d
320 × 6 × 6
ReLU
320 × 6 × 6
Upsample 1
Nearest upsample ×2
320 × 12 × 12
Conv2d 3×3 / 1
320 × 12 × 12; p=1, g=320
BatchNorm2d
320 × 12 × 12
ReLU
320 × 12 × 12
Conv2d 1×1 / 1
192 × 12 × 12; p=0, g=1
BatchNorm2d
192 × 12 × 12
ReLU
192 × 12 × 12
Upsample 2
Nearest upsample ×2
192 × 24 × 24
Conv2d 3×3 / 1
192 × 24 × 24; p=1, g=192
BatchNorm2d
192 × 24 × 24
ReLU
192 × 24 × 24
Conv2d 1×1 / 1
128 × 24 × 24; p=0, g=1
BatchNorm2d
128 × 24 × 24
ReLU
128 × 24 × 24
Upsample 3
Nearest upsample ×2
128 × 48 × 48
Conv2d 3×3 / 1
128 × 48 × 48; p=1, g=128
BatchNorm2d
128 × 48 × 48
ReLU
128 × 48 × 48
Conv2d 1×1 / 1
80 × 48 × 48; p=0, g=1
BatchNorm2d
80 × 48 × 48
ReLU
80 × 48 × 48
Upsample 4
Nearest upsample ×2
80 × 96 × 96
Conv2d 3×3 / 1
80 × 96 × 96; p=1, g=80
BatchNorm2d
80 × 96 × 96
ReLU
80 × 96 × 96
Conv2d 1×1 / 1
40 × 96 × 96; p=0, g=1
BatchNorm2d
40 × 96 × 96
ReLU
40 × 96 × 96
Refine
Conv2d 3×3 / 1
40 × 96 × 96; p=1, g=40
BatchNorm2d
40 × 96 × 96
ReLU
40 × 96 × 96
Conv2d 1×1 / 1
1 × 96 × 96; p=0, g=1
Channel attention (ECA)
Input
40 × 96 × 96
Global average pool
40 × 1 × 1
Conv2d 1×1
40 to 40, no bias
Sigmoid
40 × 1 × 1
Multiply gate and input
40 × 96 × 96
Box/point prompts select an ROI before this network. The decoder has additive projected skips; it does not concatenate encoder maps.
Source: libreyolo/models/picosam3/nn.py. Revision a4d0ecc9e17f.
libreyolo.com