YOLOv2-B
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
YOLOv2-B
Detection; 608 × 608 RGB; 80 classes; batch 1; unfused eval. Every cfg layer is shown.
LibreYOLO
YOLOv2-B
Detection; 608 × 608 RGB; 80 classes; batch 1; unfused eval. Every cfg layer is shown.
L = cfg layer index. Matching F labels continue the same tensor across columns. Full routes remain visible.
Layers 0 to 15
L0 Conv 3×3, s=1 + norm + leaky
32 × 608 × 608
L1 MaxPool2d 2×2, s=2
32 × 304 × 304
L2 Conv 3×3, s=1 + norm + leaky
64 × 304 × 304
L3 MaxPool2d 2×2, s=2
64 × 152 × 152
L4 Conv 3×3, s=1 + norm + leaky
128 × 152 × 152
L5 Conv 1×1, s=1 + norm + leaky
64 × 152 × 152
L6 Conv 3×3, s=1 + norm + leaky
128 × 152 × 152
L7 MaxPool2d 2×2, s=2
128 × 76 × 76
L8 Conv 3×3, s=1 + norm + leaky
256 × 76 × 76
L9 Conv 1×1, s=1 + norm + leaky
128 × 76 × 76
L10 Conv 3×3, s=1 + norm + leaky
256 × 76 × 76
L11 MaxPool2d 2×2, s=2
256 × 38 × 38
L12 Conv 3×3, s=1 + norm + leaky
512 × 38 × 38
L13 Conv 1×1, s=1 + norm + leaky
256 × 38 × 38
L14 Conv 3×3, s=1 + norm + leaky
512 × 38 × 38
L15 Conv 1×1, s=1 + norm + leaky
256 × 38 × 38
Input: 3 × 608 × 608
F15
Layers 16 to 31
L16 Conv 3×3, s=1 + norm + leaky
512 × 38 × 38
L17 MaxPool2d 2×2, s=2
512 × 19 × 19
L18 Conv 3×3, s=1 + norm + leaky
1024 × 19 × 19
L19 Conv 1×1, s=1 + norm + leaky
512 × 19 × 19
L20 Conv 3×3, s=1 + norm + leaky
1024 × 19 × 19
L21 Conv 1×1, s=1 + norm + leaky
512 × 19 × 19
L22 Conv 3×3, s=1 + norm + leaky
1024 × 19 × 19
L23 Conv 3×3, s=1 + norm + leaky
1024 × 19 × 19
L24 Conv 3×3, s=1 + norm + leaky
1024 × 19 × 19
L25 Route
512 × 38 × 38
L26 Conv 1×1, s=1 + norm + leaky
64 × 38 × 38
L27 Reorg, stride 2
256 × 19 × 19
L28 Concat
1280 × 19 × 19
L29 Conv 3×3, s=1 + norm + leaky
1024 × 19 × 19
L30 Conv 1×1, s=1
425 × 19 × 19
L31 Raw region head
425 × 19 × 19
F15
Convolution block
Conv2d
k, stride and channels from layer label
Darknet normalization
Only layers labeled + norm
Activation
leaky: LeakyReLU(0.1); linear: identity
Norm layers: bias=False. Without norm: bias=True.
Mish uses its separate definition when present.
Normalization and pooling
Subtract running mean
x - mean
Divide by standard deviation
sqrt(running variance) + 0.000001
Multiply scale, add bias
Learned channel-wise affine transform
MaxPool2d: p = floor((k - 1) / 2).
k=2, s=1: pad right/bottom with negative infinity.
That pool preserves its input spatial size.
Reorg (Darknet channel ordering)
Reshape + transpose axes 3, 4
1 × 64 × 19 × 2 × 19 × 2
Reshape + transpose axes 2, 3
1 × 64 × 361 × 4
Reshape + transpose axes 1, 2
1 × 64 × 4 × 19 × 19
Reshape contiguous result
1 × 256 × 19 × 19
Detection decoding (postprocessing)
Raw head outputs: 425 × 19 × 19
Raw prediction tensor
Field selection is explicit in the three branches
Select xy; sigmoid center offsets
Grid + offsets; multiply by per-head stride
Select wh; exponentiate logits
Multiply by anchor widths and heights
Select objectness and class logits
Sigmoid objectness × softmax classes
Stack cx, cy, width, height
Convert to corner coordinates
Confidence filter + class NMS
Then invert resize to the original image
Anchors by raw head (width, height): [(0.57273, 0.677385), (1.87446, 2.06253), (3.33843, 5.47434), (7.88282, 3.52778), (9.77052, 9.16828)]
Strides: 32. Region anchors are in grid cells, so width/height also multiply by stride.
Center scale_x_y per head: 1.0. Offset = sigmoid(raw) × scale - (scale - 1)/2.
Source: darknet/cfgs/yolov2.cfg; darknet/net.py; darknet/blocks.py. Revision a4d0ecc9e17f.
libreyolo.com