PicoSAM3 pico

Click a block to read its description, or select it with Tab and Enter.

PicoSAM3 picoROI mask segmentation, 96 × 96 input. One mask logit map. Tensor sizes exclude batch.LibreYOLOPicoSAM3 picoROI mask segmentation, 96 × 96 input. One mask logit map. Tensor sizes exclude batch.Encoder and decoderROI image3 × 96 × 96Encoder stage 148 × 96 × 96Conv2d 3×3 / 248 × 48 × 48, p=1Skip Conv2d 1×148 to 40, bias=TrueUpsample block 440 × 96 × 96+Encoder stage 296 × 48 × 48Conv2d 3×3 / 296 × 24 × 24, p=1Skip Conv2d 1×196 to 80, bias=TrueUpsample block 380 × 48 × 48+Encoder stage 3160 × 24 × 24Conv2d 3×3 / 2160 × 12 × 12, p=1Skip Conv2d 1×1160 to 128, bias=TrueUpsample block 2128 × 24 × 24+Encoder stage 4256 × 12 × 12Conv2d 3×3 / 2256 × 6 × 6, p=1Skip Conv2d 1×1256 to 192, bias=TrueUpsample block 1192 × 12 × 12+Bottleneck320 × 6 × 6Channel attention (ECA)40 × 96 × 96Refine mask logits1 × 96 × 96Sigmoid and ROI resizeMask pasted into original imageEncoder 1Conv2d 3×3 / 13 × 96 × 96; p=1, g=3BatchNorm2d3 × 96 × 96ReLU3 × 96 × 96Conv2d 1×1 / 148 × 96 × 96; p=0, g=1BatchNorm2d48 × 96 × 96ReLU48 × 96 × 96Encoder 2Conv2d 3×3 / 148 × 48 × 48; p=1, g=48BatchNorm2d48 × 48 × 48ReLU48 × 48 × 48Conv2d 1×1 / 196 × 48 × 48; p=0, g=1BatchNorm2d96 × 48 × 48ReLU96 × 48 × 48Encoder 3Conv2d 3×3 / 196 × 24 × 24; p=1, g=96BatchNorm2d96 × 24 × 24ReLU96 × 24 × 24Conv2d 1×1 / 1160 × 24 × 24; p=0, g=1BatchNorm2d160 × 24 × 24ReLU160 × 24 × 24Encoder 4Conv2d 3×3 / 1160 × 12 × 12; p=1, g=160BatchNorm2d160 × 12 × 12ReLU160 × 12 × 12Conv2d 1×1 / 1256 × 12 × 12; p=0, g=1BatchNorm2d256 × 12 × 12ReLU256 × 12 × 12BottleneckConv2d 3×3 / 1256 × 6 × 6; p=1, g=256BatchNorm2d256 × 6 × 6ReLU256 × 6 × 6Conv2d 1×1 / 1320 × 6 × 6; p=0, g=1BatchNorm2d320 × 6 × 6ReLU320 × 6 × 6Conv2d 3×3 / 1320 × 6 × 6; p=2, g=320, d=2BatchNorm2d320 × 6 × 6ReLU320 × 6 × 6Conv2d 1×1 / 1320 × 6 × 6; p=0, g=1BatchNorm2d320 × 6 × 6ReLU320 × 6 × 6Upsample 1Nearest upsample ×2320 × 12 × 12Conv2d 3×3 / 1320 × 12 × 12; p=1, g=320BatchNorm2d320 × 12 × 12ReLU320 × 12 × 12Conv2d 1×1 / 1192 × 12 × 12; p=0, g=1BatchNorm2d192 × 12 × 12ReLU192 × 12 × 12Upsample 2Nearest upsample ×2192 × 24 × 24Conv2d 3×3 / 1192 × 24 × 24; p=1, g=192BatchNorm2d192 × 24 × 24ReLU192 × 24 × 24Conv2d 1×1 / 1128 × 24 × 24; p=0, g=1BatchNorm2d128 × 24 × 24ReLU128 × 24 × 24Upsample 3Nearest upsample ×2128 × 48 × 48Conv2d 3×3 / 1128 × 48 × 48; p=1, g=128BatchNorm2d128 × 48 × 48ReLU128 × 48 × 48Conv2d 1×1 / 180 × 48 × 48; p=0, g=1BatchNorm2d80 × 48 × 48ReLU80 × 48 × 48Upsample 4Nearest upsample ×280 × 96 × 96Conv2d 3×3 / 180 × 96 × 96; p=1, g=80BatchNorm2d80 × 96 × 96ReLU80 × 96 × 96Conv2d 1×1 / 140 × 96 × 96; p=0, g=1BatchNorm2d40 × 96 × 96ReLU40 × 96 × 96RefineConv2d 3×3 / 140 × 96 × 96; p=1, g=40BatchNorm2d40 × 96 × 96ReLU40 × 96 × 96Conv2d 1×1 / 11 × 96 × 96; p=0, g=1Channel attention (ECA)Input40 × 96 × 96Global average pool40 × 1 × 1Conv2d 1×140 to 40, no biasSigmoid40 × 1 × 1Multiply gate and input40 × 96 × 96Box/point prompts select an ROI before this network. The decoder has additive projected skips; it does not concatenate encoder maps.Source: libreyolo/models/picosam3/nn.py. Revision a4d0ecc9e17f.libreyolo.com