SSD300-VGG16
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
SSD300-VGG16
Detection; 300 × 300 RGB; batch 1. COCO: 91 head slots, including background and unused IDs, mapped to 80 classes.
LibreYOLO
SSD300-VGG16
Detection; 300 × 300 RGB; batch 1. COCO: 91 head slots, including background and unused IDs, mapped to 80 classes.
VGG16 and SSD feature extractor
Input
3 × 300 × 300
Conv-ReLU 3×3 repeated 2 times
64 × 300 × 300
MaxPool2d 2×2, s=2
64 × 150 × 150
Conv-ReLU 3×3 repeated 2 times
128 × 150 × 150
MaxPool2d 2×2, s=2
128 × 75 × 75
Conv-ReLU 3×3 repeated 3 times
256 × 75 × 75
MaxPool2d 2×2, s=2, ceil=True
256 × 38 × 38
Conv-ReLU 3×3 repeated 3 times
512 × 38 × 38; unscaled continuation
MaxPool2d 2×2, s=2
512 × 19 × 19
Conv-ReLU 3×3 repeated 3 times
512 × 19 × 19
MaxPool2d 3×3, s=1, p=1
512 × 19 × 19
Conv-ReLU 3×3, dilation=6, p=6
1024 × 19 × 19
Conv-ReLU 1×1
1024 × 19 × 19 (F1)
Extra block
512 × 10 × 10 (F2)
Extra block
256 × 5 × 5 (F3)
Extra block
256 × 3 × 3 (F4)
Extra block
256 × 1 × 1 (F5)
All repeated convs use p=1 and s=1 unless marked.
F0 feature normalization
Unscaled conv4_3
512 × 38 × 38
L2 normalization across channels
Divide by max(L2 norm, 0.000000000001)
Multiply learned channel scale
512 weights, initialized to 20; F0 = 512 × 38 × 38
Only F0 is normalized. The unscaled conv4_3 feature continues through pool4 to the later feature blocks.
Six independent box and class heads
F0
512 × 38 × 38
Conv2d 3×3, p=1
16 output channels; bias=True
Conv2d 3×3, p=1
364 output channels; bias=True
Reshape and permute
1 × 5,776 × 4
Reshape and permute
1 × 5,776 × 91
B0
C0
F1
1024 × 19 × 19
Conv2d 3×3, p=1
24 output channels; bias=True
Conv2d 3×3, p=1
546 output channels; bias=True
Reshape and permute
1 × 2,166 × 4
Reshape and permute
1 × 2,166 × 91
B1
C1
F2
512 × 10 × 10
Conv2d 3×3, p=1
24 output channels; bias=True
Conv2d 3×3, p=1
546 output channels; bias=True
Reshape and permute
1 × 600 × 4
Reshape and permute
1 × 600 × 91
B2
C2
F3
256 × 5 × 5
Conv2d 3×3, p=1
24 output channels; bias=True
Conv2d 3×3, p=1
546 output channels; bias=True
Reshape and permute
1 × 150 × 4
Reshape and permute
1 × 150 × 91
B3
C3
F4
256 × 3 × 3
Conv2d 3×3, p=1
16 output channels; bias=True
Conv2d 3×3, p=1
364 output channels; bias=True
Reshape and permute
1 × 36 × 4
Reshape and permute
1 × 36 × 91
B4
C4
F5
256 × 1 × 1
Conv2d 3×3, p=1
16 output channels; bias=True
Conv2d 3×3, p=1
364 output channels; bias=True
Reshape and permute
1 × 4 × 4
Reshape and permute
1 × 4 × 91
B5
C5
Concat box levels
1 × 8,732 × 4
Concat class levels
1 × 8,732 × 91
B0
C0
B1
C1
B2
C2
B3
C3
B4
C4
B5
C5
Concat box rows across F0...F5: 1 × 8,732 × 4. Concat class rows: 1 × 8,732 × 91.
Anchor-row order is spatial location, then anchor; every level has its own prediction convolutions.
Conv-ReLU
Conv2d
Bias=True; k/s/p from stage
ReLU
No BatchNorm
Extra block (fully resolved)
Each line is Conv2d 1×1 + ReLU, then Conv2d 3×3 + ReLU.
F2: 1024 input; 256 hidden; 512 output; second conv s=2, p=1.
F3: 512 input; 128 hidden; 256 output; second conv s=2, p=1.
F4: 256 input; 128 hidden; 256 output; second conv s=1, p=0.
F5: 256 input; 128 hidden; 256 output; second conv s=1, p=0.
Extra-block internals and SSD decoding
Conv2d 1×1
Hidden widths listed above
ReLU
Conv2d 3×3
Output/stride/padding listed above
ReLU
Box offsets: 8,732 × 4
Divide center deltas by 10; size deltas by 5
Generate 8,732 default boxes
Steps: 8, 16, 32, 64, 100, 300; centers have +0.5 offset
Class logits: 8,732 × 91
Softmax over 91 slots
Apply deltas to default-box centers and sizes
Size exponent is clamped at log(1000/16); exp sizes; convert xyxy
Drop background; map sparse COCO IDs
Threshold; top-400 per class
Class-wise NMS; retain at most 200; invert image resize
Final boxes, scores and 80-class IDs
Default-box scales: 0.07, 0.15, 0.33, 0.51, 0.69, 0.87; next scale 1.05. Aspect ratios per level: [2], [2,3], [2,3], [2,3], [2], [2].
Each level includes its scale square and geometric-mean square, plus reciprocal aspect-ratio pairs. Sizes are clipped to [0,1] before pixel scaling.
Source: models/ssd/nn.py; postprocess/ssd.py. Revision a4d0ecc9e17f.
libreyolo.com