RetinaNet R50
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
RetinaNet R50
Detection; 800 × 800 RGB canvas; batch 1; unfused eval. 91 class-head slots map to 80 COCO classes.
LibreYOLO
RetinaNet R50
Detection; 800 × 800 RGB canvas; batch 1; unfused eval. 91 class-head slots map to 80 COCO classes.
ResNet-50 backbone
Input
3 × 800 × 800
Conv2d 7×7, s=2, p=3
64 × 400 × 400; no bias
FrozenBatchNorm2d + ReLU
64 × 400 × 400; eps=0
MaxPool2d 3×3, s=2, p=1
64 × 200 × 200
C2: 1 projection + 2 identity blocks
256 × 200 × 200; hidden 64
C3: 1 projection + 3 identity blocks
512 × 100 × 100; hidden 128
C4: 1 projection + 5 identity blocks
1024 × 50 × 50; hidden 256
C5: 1 projection + 2 identity blocks
2048 × 25 × 25; hidden 512
Projection strides: C2=1, C3/C4/C5=2.
The stride is in the bottleneck 3×3 convolution.
Feature pyramid (FPN)
C5: Conv2d 1×1
2048 input; 256 × 25 × 25
P5: Conv2d 3×3, p=1
256 × 25 × 25
C4: Conv2d 1×1
1024 input; 256 × 50 × 50
P4: Conv2d 3×3, p=1
256 × 50 × 50
+
C3: Conv2d 1×1
512 input; 256 × 100 × 100
P3: Conv2d 3×3, p=1
256 × 100 × 100
+
Nearest resize
256 × 50 × 50
Nearest resize
256 × 100 × 100
P5 continuation
256 × 25 × 25
Conv2d 3×3, s=2, p=1
P6: 256 × 13 × 13
ReLU
P6 activation
Conv2d 3×3, s=2
P7: 256 × 7 × 7; p=1
Lateral sums feed upsampling before the output 3×3 convolutions.
RetinaNet head
Apply the same tower weights to P3, P4, P5, P6 and P7.
One pyramid feature
256 channels; grid 100 / 50 / 25 / 13 / 7
Head unit repeated 4 times
Conv3×3 + ReLU; 256 ch
Head unit repeated 4 times
Conv3×3 + ReLU; 256 ch
Conv2d 3×3, p=1
819 logits/location; bias=True
Reshape anchor/location rows
1 × 120,087 × 91
Conv2d 3×3, p=1
36 deltas/location; bias=True
Reshape anchor rows
1 × 120,087 × 4
Class and regression towers have separate parameters.
Intermediate conv bias: disabled with GroupNorm; enabled without norm.
Pyramid grid predictions per location rows
100 × 100
9
90,000
50 × 50
9
22,500
25 × 25
9
5,625
13 × 13
9
1,521
7 × 7
9
441
RetinaNet decodes boxes and sigmoid scores inside forward().
Projection bottleneck
Stage order: C2 / C3 / C4 / C5.
Input
64 / 256 / 512 / 1024 ch
Conv2d 1×1
64 / 128 / 256 / 512 output channels
FrozenBatchNorm2d
64 / 128 / 256 / 512 channels
ReLU
Conv2d 3×3, p=1
s=1/2/2/2
FrozenBatchNorm2d
64 / 128 / 256 / 512 channels
ReLU
Conv2d 1×1
256 / 512 / 1024 / 2048 output channels
FrozenBatchNorm2d
No activation before addition
+
ReLU
Block output
Conv2d 1×1
256 / 512 / 1024 / 2048 ch
FrozenBatchNorm2d
Stride 1 / 2 / 2 / 2
Identity bottleneck
Stage order: C2 / C3 / C4 / C5.
Input
256 / 512 / 1024 / 2048 ch
Conv2d 1×1
64 / 128 / 256 / 512 output channels
FrozenBatchNorm2d
64 / 128 / 256 / 512 channels
ReLU
Conv2d 3×3, p=1
s=1
FrozenBatchNorm2d
64 / 128 / 256 / 512 channels
ReLU
Conv2d 1×1
256 / 512 / 1024 / 2048 output channels
FrozenBatchNorm2d
No activation before addition
+
ReLU
Block output
identity
Normalization and head unit
FrozenBatchNorm2d
eps=0; learned scale and bias
ReLU
For stem and first two bottleneck convolutions
Frozen normalization uses stored mean/variance;
BatchNorm in eval also uses its running statistics.
Affine transform: (x - mean) / sqrt(variance + eps),
then multiply scale and add bias.
Head Conv2d 3×3, p=1
256 input/output channels
Optional GroupNorm
32 groups; eps=0.00001
ReLU
This head unit repeats four times in each tower
RetinaNet r50 has no head normalization.
RetinaNet r50v2 and FCOS use GroupNorm.
Location/anchor decoding and selection
Nine anchors per location
Three scales × ratios 0.5, 1, 2
Decode center/size deltas
Weights all 1; exp sizes clamped at log(1000/16)
Sigmoid classes; map COCO slots
Network output: 1 × 120,087 × 84
Confidence filtering and class NMS
Postprocessing follows network forward
Actual 800-pixel canvas grid spacing is floor(800/grid): 8,16,32,61,114. Anchor base sizes remain separately defined.
Base anchor sizes per level: (32,40,50), (64,80,101), (128,161,203), (256,322,406), (512,645,812).
Random-weight CPU checks cover backbone, FPN, raw heads and model outputs. No pretrained weights were downloaded.
Source: models/retinanet/nn.py; postprocess/retinanet.py. Revision a4d0ecc9e17f.
libreyolo.com