YOLO7-B
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
YOLO7-B
Detection; 640 × 640 RGB; 80 classes; batch 1; unfused native eval. Layer numbers preserve the MIT v7.yaml graph.
LibreYOLO
YOLO7-B
Detection; 640 × 640 RGB; 80 classes; batch 1; unfused native eval. Layer numbers preserve the MIT v7.yaml graph.
Matching F labels continue a tensor across columns. All YAML branches are explicit; block internals follow the in-tree implementation.
Layers 1 to 27
L1 Conv 3×3, s=1
32 ×640 ×640
L2 Conv 3×3, s=2
64 ×320 ×320
L3 Conv 3×3, s=1
64 ×320 ×320
L4 Conv 3×3, s=2
128 ×160 ×160
L5 Conv 1×1, s=1
64 ×160 ×160
L6 Conv 1×1, s=1
64 ×160 ×160
L7 Conv 3×3, s=1
64 ×160 ×160
L8 Conv 3×3, s=1
64 ×160 ×160
L9 Conv 3×3, s=1
64 ×160 ×160
L10 Conv 3×3, s=1
64 ×160 ×160
L11 Concat
256 ×160 ×160
L12 Conv 1×1, s=1
256 ×160 ×160
L13 MaxPool 2×2, s=2, p=0
256 ×80 ×80
L14 Conv 1×1, s=1
128 ×80 ×80
L15 Conv 1×1, s=1
128 ×160 ×160
L16 Conv 3×3, s=2
128 ×80 ×80
L17 Concat
256 ×80 ×80
L18 Conv 1×1, s=1
128 ×80 ×80
L19 Conv 1×1, s=1
128 ×80 ×80
L20 Conv 3×3, s=1
128 ×80 ×80
L21 Conv 3×3, s=1
128 ×80 ×80
L22 Conv 3×3, s=1
128 ×80 ×80
L23 Conv 3×3, s=1
128 ×80 ×80
L24 Concat (B3)
512 ×80 ×80
L25 Conv 1×1, s=1
512 ×80 ×80
L26 MaxPool 2×2, s=2, p=0
512 ×40 ×40
L27 Conv 1×1, s=1
256 ×40 ×40
Input: 3 × 640 × 640
F24
F25
F27
Layers 28 to 54
L28 Conv 1×1, s=1
256 ×80 ×80
L29 Conv 3×3, s=2
256 ×40 ×40
L30 Concat
512 ×40 ×40
L31 Conv 1×1, s=1
256 ×40 ×40
L32 Conv 1×1, s=1
256 ×40 ×40
L33 Conv 3×3, s=1
256 ×40 ×40
L34 Conv 3×3, s=1
256 ×40 ×40
L35 Conv 3×3, s=1
256 ×40 ×40
L36 Conv 3×3, s=1
256 ×40 ×40
L37 Concat
1024 ×40 ×40
L38 Conv 1×1, s=1 (B4)
1024 ×40 ×40
L39 MaxPool 2×2, s=2, p=0
1024 ×20 ×20
L40 Conv 1×1, s=1
512 ×20 ×20
L41 Conv 1×1, s=1
512 ×40 ×40
L42 Conv 3×3, s=2
512 ×20 ×20
L43 Concat
1024 ×20 ×20
L44 Conv 1×1, s=1
256 ×20 ×20
L45 Conv 1×1, s=1
256 ×20 ×20
L46 Conv 3×3, s=1
256 ×20 ×20
L47 Conv 3×3, s=1
256 ×20 ×20
L48 Conv 3×3, s=1
256 ×20 ×20
L49 Conv 3×3, s=1
256 ×20 ×20
L50 Concat
1024 ×20 ×20
L51 Conv 1×1, s=1 (B5)
1024 ×20 ×20
L52 SPPCSPConv (N3)
512 ×20 ×20
L53 Conv 1×1, s=1
256 ×20 ×20
L54 Nearest upsample ×2
256 ×40 ×40
F25
F27
F38
F52
F54
Layers 55 to 81
L55 Conv 1×1, s=1
256 ×40 ×40
L56 Concat
512 ×40 ×40
L57 Conv 1×1, s=1
256 ×40 ×40
L58 Conv 1×1, s=1
256 ×40 ×40
L59 Conv 3×3, s=1
128 ×40 ×40
L60 Conv 3×3, s=1
128 ×40 ×40
L61 Conv 3×3, s=1
128 ×40 ×40
L62 Conv 3×3, s=1
128 ×40 ×40
L63 Concat
1024 ×40 ×40
L64 Conv 1×1, s=1 (N2)
256 ×40 ×40
L65 Conv 1×1, s=1
128 ×40 ×40
L66 Nearest upsample ×2
128 ×80 ×80
L67 Conv 1×1, s=1
128 ×80 ×80
L68 Concat
256 ×80 ×80
L69 Conv 1×1, s=1
128 ×80 ×80
L70 Conv 1×1, s=1
128 ×80 ×80
L71 Conv 3×3, s=1
64 ×80 ×80
L72 Conv 3×3, s=1
64 ×80 ×80
L73 Conv 3×3, s=1
64 ×80 ×80
L74 Conv 3×3, s=1
64 ×80 ×80
L75 Concat
512 ×80 ×80
L76 Conv 1×1, s=1 (P3)
128 ×80 ×80
L77 MaxPool 2×2, s=2, p=0
128 ×40 ×40
L78 Conv 1×1, s=1
128 ×40 ×40
L79 Conv 1×1, s=1
128 ×80 ×80
L80 Conv 3×3, s=2
128 ×40 ×40
L81 Concat
512 ×40 ×40
F38
F54
F24
F76
F81
Layers 82 to 106
L82 Conv 1×1, s=1
256 ×40 ×40
L83 Conv 1×1, s=1
256 ×40 ×40
L84 Conv 3×3, s=1
128 ×40 ×40
L85 Conv 3×3, s=1
128 ×40 ×40
L86 Conv 3×3, s=1
128 ×40 ×40
L87 Conv 3×3, s=1
128 ×40 ×40
L88 Concat
1024 ×40 ×40
L89 Conv 1×1, s=1 (P4)
256 ×40 ×40
L90 MaxPool 2×2, s=2, p=0
256 ×20 ×20
L91 Conv 1×1, s=1
256 ×20 ×20
L92 Conv 1×1, s=1
256 ×40 ×40
L93 Conv 3×3, s=2
256 ×20 ×20
L94 Concat
1024 ×20 ×20
L95 Conv 1×1, s=1
512 ×20 ×20
L96 Conv 1×1, s=1
512 ×20 ×20
L97 Conv 3×3, s=1
256 ×20 ×20
L98 Conv 3×3, s=1
256 ×20 ×20
L99 Conv 3×3, s=1
256 ×20 ×20
L100 Conv 3×3, s=1
256 ×20 ×20
L101 Concat
2048 ×20 ×20
L102 Conv 1×1, s=1 (P5)
512 ×20 ×20
L103 RepConv
256 ×80 ×80
L104 RepConv
512 ×40 ×40
L105 RepConv
1024 ×20 ×20
L106 MultiheadDetection (Main)
255 ×80 ×80 / 255 ×40 ×40 / 255 ×20 ×20
F81
F81
F52
F76
Conv and RepConv
Conv2d
No bias; p=(k-1)/2; groups 1
BatchNorm2d
eps=.001; momentum=.03
SiLU
Disabled inside the two RepConv branches
RepConv input
128 /256 /512 channels
Conv2d 3×3
256 /512 /1024 output channels
BatchNorm2d
eps=.001; no branch activation
Conv2d 1×1
256 /512 /1024 output channels
BatchNorm2d
eps=.001; no branch activation
+
SiLU
No identity BN branch
SPPCSPConv
Input
1024 ×20 ×20
Conv 1×1
512 channels
Conv 3×3
512 channels
Conv 1×1
512 channels
MaxPool 5×5
s=1, p=2;512 channels
MaxPool 9×9
s=1, p=4;512 channels
MaxPool 13×13
s=1, p=6;512 channels
Concat four sequential taps
2048 channels
Conv 1×1
512 channels
Conv 3×3
512 channels
Concat main and short branches
1024 channels
Conv 1×1
512 ×20 ×20
Conv 1×1
1024 input;512 output
The in-tree pools are sequential, with kernels 5 then 9 then 13.
Implicit detection head and decoding
Each scale separately
256×80² /512×40² /1024×20²
Add learned per-channel ImplicitA
256 /512 /1024 bias parameters
Conv2d1×1
255 output channels; bias=True
Multiply learned ImplicitM
255 scale parameters per head
Raw predictions
255 ×80²;255 ×40²;255 ×20²
Reshape 3 anchors ×85 values; sigmoid
Applied to all five box/objectness fields and 80 classes
Center coordinates
(2×sigmoid(xy)-.5+grid) ×stride
Width and height
(2×sigmoid(wh))² ×anchor size
Class confidence
sigmoid(objectness) ×sigmoid(class)
Strides 8/16/32. Per-scale anchor (w,h) pairs:
[(12, 16), (19, 36), (40, 28)]
[(36, 75), (76, 55), (72, 146)]
[(142, 110), (192, 243), (459, 401)]
Flatten 25,200 anchor rows; convert to xyxy; threshold and class NMS.
No pretrained weights were used for the CPU shape check.
Source: models/yolo7/v7.yaml; models/yolo7/blocks.py. Revision a4d0ecc9e17f.
libreyolo.com