SSD300-VGG16

Click a block to read its description, or select it with Tab and Enter.

SSD300-VGG16Detection; 300 × 300 RGB; batch 1. COCO: 91 head slots, including background and unused IDs, mapped to 80 classes.LibreYOLOSSD300-VGG16Detection; 300 × 300 RGB; batch 1. COCO: 91 head slots, including background and unused IDs, mapped to 80 classes.VGG16 and SSD feature extractorInput3 × 300 × 300Conv-ReLU 3×3 repeated 2 times64 × 300 × 300MaxPool2d 2×2, s=264 × 150 × 150Conv-ReLU 3×3 repeated 2 times128 × 150 × 150MaxPool2d 2×2, s=2128 × 75 × 75Conv-ReLU 3×3 repeated 3 times256 × 75 × 75MaxPool2d 2×2, s=2, ceil=True256 × 38 × 38Conv-ReLU 3×3 repeated 3 times512 × 38 × 38; unscaled continuationMaxPool2d 2×2, s=2512 × 19 × 19Conv-ReLU 3×3 repeated 3 times512 × 19 × 19MaxPool2d 3×3, s=1, p=1512 × 19 × 19Conv-ReLU 3×3, dilation=6, p=61024 × 19 × 19Conv-ReLU 1×11024 × 19 × 19 (F1)Extra block512 × 10 × 10 (F2)Extra block256 × 5 × 5 (F3)Extra block256 × 3 × 3 (F4)Extra block256 × 1 × 1 (F5)All repeated convs use p=1 and s=1 unless marked.F0 feature normalizationUnscaled conv4_3512 × 38 × 38L2 normalization across channelsDivide by max(L2 norm, 0.000000000001)Multiply learned channel scale512 weights, initialized to 20; F0 = 512 × 38 × 38Only F0 is normalized. The unscaled conv4_3 feature continues through pool4 to the later feature blocks.Six independent box and class headsF0512 × 38 × 38Conv2d 3×3, p=116 output channels; bias=TrueConv2d 3×3, p=1364 output channels; bias=TrueReshape and permute1 × 5,776 × 4Reshape and permute1 × 5,776 × 91B0C0F11024 × 19 × 19Conv2d 3×3, p=124 output channels; bias=TrueConv2d 3×3, p=1546 output channels; bias=TrueReshape and permute1 × 2,166 × 4Reshape and permute1 × 2,166 × 91B1C1F2512 × 10 × 10Conv2d 3×3, p=124 output channels; bias=TrueConv2d 3×3, p=1546 output channels; bias=TrueReshape and permute1 × 600 × 4Reshape and permute1 × 600 × 91B2C2F3256 × 5 × 5Conv2d 3×3, p=124 output channels; bias=TrueConv2d 3×3, p=1546 output channels; bias=TrueReshape and permute1 × 150 × 4Reshape and permute1 × 150 × 91B3C3F4256 × 3 × 3Conv2d 3×3, p=116 output channels; bias=TrueConv2d 3×3, p=1364 output channels; bias=TrueReshape and permute1 × 36 × 4Reshape and permute1 × 36 × 91B4C4F5256 × 1 × 1Conv2d 3×3, p=116 output channels; bias=TrueConv2d 3×3, p=1364 output channels; bias=TrueReshape and permute1 × 4 × 4Reshape and permute1 × 4 × 91B5C5Concat box levels1 × 8,732 × 4Concat class levels1 × 8,732 × 91B0C0B1C1B2C2B3C3B4C4B5C5Concat box rows across F0...F5: 1 × 8,732 × 4. Concat class rows: 1 × 8,732 × 91.Anchor-row order is spatial location, then anchor; every level has its own prediction convolutions.Conv-ReLUConv2dBias=True; k/s/p from stageReLUNo BatchNormExtra block (fully resolved)Each line is Conv2d 1×1 + ReLU, then Conv2d 3×3 + ReLU.F2: 1024 input; 256 hidden; 512 output; second conv s=2, p=1.F3: 512 input; 128 hidden; 256 output; second conv s=2, p=1.F4: 256 input; 128 hidden; 256 output; second conv s=1, p=0.F5: 256 input; 128 hidden; 256 output; second conv s=1, p=0.Extra-block internals and SSD decodingConv2d 1×1Hidden widths listed aboveReLUConv2d 3×3Output/stride/padding listed aboveReLUBox offsets: 8,732 × 4Divide center deltas by 10; size deltas by 5Generate 8,732 default boxesSteps: 8, 16, 32, 64, 100, 300; centers have +0.5 offsetClass logits: 8,732 × 91Softmax over 91 slotsApply deltas to default-box centers and sizesSize exponent is clamped at log(1000/16); exp sizes; convert xyxyDrop background; map sparse COCO IDsThreshold; top-400 per classClass-wise NMS; retain at most 200; invert image resizeFinal boxes, scores and 80-class IDsDefault-box scales: 0.07, 0.15, 0.33, 0.51, 0.69, 0.87; next scale 1.05. Aspect ratios per level: [2], [2,3], [2,3], [2,3], [2], [2].Each level includes its scale square and geometric-mean square, plus reciprocal aspect-ratio pairs. Sizes are clipped to [0,1] before pixel scaling.Source: models/ssd/nn.py; postprocess/ssd.py. Revision a4d0ecc9e17f.libreyolo.com