YOLOv1-B
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
YOLOv1-B
Detection; 448 × 448 RGB; 20 classes; batch 1; unfused eval. Every cfg layer is shown.
LibreYOLO
YOLOv1-B
Detection; 448 × 448 RGB; 20 classes; batch 1; unfused eval. Every cfg layer is shown.
L = cfg layer index. Matching F labels continue the same tensor across columns. Full routes remain visible.
Layers 0 to 15
L0 Conv 7×7, s=2 + norm + leaky
64 × 224 × 224
L1 MaxPool2d 2×2, s=2
64 × 112 × 112
L2 Conv 3×3, s=1 + norm + leaky
192 × 112 × 112
L3 MaxPool2d 2×2, s=2
192 × 56 × 56
L4 Conv 1×1, s=1 + norm + leaky
128 × 56 × 56
L5 Conv 3×3, s=1 + norm + leaky
256 × 56 × 56
L6 Conv 1×1, s=1 + norm + leaky
256 × 56 × 56
L7 Conv 3×3, s=1 + norm + leaky
512 × 56 × 56
L8 MaxPool2d 2×2, s=2
512 × 28 × 28
L9 Conv 1×1, s=1 + norm + leaky
256 × 28 × 28
L10 Conv 3×3, s=1 + norm + leaky
512 × 28 × 28
L11 Conv 1×1, s=1 + norm + leaky
256 × 28 × 28
L12 Conv 3×3, s=1 + norm + leaky
512 × 28 × 28
L13 Conv 1×1, s=1 + norm + leaky
256 × 28 × 28
L14 Conv 3×3, s=1 + norm + leaky
512 × 28 × 28
L15 Conv 1×1, s=1 + norm + leaky
256 × 28 × 28
Input: 3 × 448 × 448
F15
Layers 16 to 31
L16 Conv 3×3, s=1 + norm + leaky
512 × 28 × 28
L17 Conv 1×1, s=1 + norm + leaky
512 × 28 × 28
L18 Conv 3×3, s=1 + norm + leaky
1024 × 28 × 28
L19 MaxPool2d 2×2, s=2
1024 × 14 × 14
L20 Conv 1×1, s=1 + norm + leaky
512 × 14 × 14
L21 Conv 3×3, s=1 + norm + leaky
1024 × 14 × 14
L22 Conv 1×1, s=1 + norm + leaky
512 × 14 × 14
L23 Conv 3×3, s=1 + norm + leaky
1024 × 14 × 14
L24 Conv 3×3, s=1 + norm + leaky
1024 × 14 × 14
L25 Conv 3×3, s=2 + norm + leaky
1024 × 7 × 7
L26 Conv 3×3, s=1 + norm + leaky
1024 × 7 × 7
L27 Conv 3×3, s=1 + norm + leaky
1024 × 7 × 7
L28 Local 3×3, leaky
256 × 7 × 7
L29 Dropout (identity in eval)
256 × 7 × 7
L30 Flatten + Linear
1470
L31 Raw detection head
1470
F15
Convolution block
Conv2d
k, stride and channels from layer label
Darknet normalization
Only layers labeled + norm
Activation
leaky: LeakyReLU(0.1); linear: identity
Norm layers: bias=False. Without norm: bias=True.
Mish uses its separate definition when present.
Normalization and pooling
Subtract running mean
x - mean
Divide by standard deviation
sqrt(running variance) + 0.000001
Multiply scale, add bias
Learned channel-wise affine transform
MaxPool2d: p = floor((k - 1) / 2).
k=2, s=1: pad right/bottom with negative infinity.
That pool preserves its input spatial size.
Fully connected head
Flatten C, H, W
12,544 features; channel-major
Linear
12,544 inputs; 1,470 outputs; bias
Activation
Identity for the final prediction vector
Locally connected layer
Unfold image patches
kernel 3; padding 1; stride 1
Per-location matrix multiply
49 separate banks; 256 × 9216 each
Add per-location bias
256 × 49 biases
Reshape + LeakyReLU(0.1)
256 × 7 × 7
Detection decoding (postprocessing)
Raw head outputs: 1470
Raw prediction tensor
Field selection is explicit in the three branches
Select xy; add grid; divide by 7
Centers scaled to 448 pixels
Select wh; square dimensions
Scale dimensions to 448 pixels
Select confidence and class values
Multiply; 20 scores for each of 98 boxes
Stack cx, cy, width, height
Convert to corner coordinates
Confidence filter + class NMS
Then invert resize to the original image
VOC 20 classes. No anchors. Shared class values per 7 × 7 cell; 2 boxes per cell. Stretch preprocessing.
Source: darknet/cfgs/yolov1.cfg; darknet/net.py; darknet/blocks.py. Revision a4d0ecc9e17f.
libreyolo.com