YOLO9-T
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
YOLO9-T
Detection; 640 × 640 RGB; 80 classes; batch 1. Unfused PyTorch eval; tensor sizes exclude batch.
LibreYOLO
YOLO9-T
Detection; 640 × 640 RGB; 80 classes; batch 1. Unfused PyTorch eval; tensor sizes exclude batch.
Matching tensor labels continue branches between panels; all block definitions and head operations are visible.
Backbone
Input
3 × 640 × 640
Conv 3×3, s=2
16 × 320 × 320
Conv 3×3, s=2
32 × 160 × 160
ELAN (B2)
32 × 160 × 160
AConv
64 × 80 × 80
RepNCSPELAN (B3)
64 × 80 × 80; part=64
AConv
96 × 40 × 40
RepNCSPELAN (B4)
96 × 40 × 40; part=96
AConv
128 × 20 × 20
RepNCSPELAN (B5)
128 × 20 × 20; part=128
SPPELAN
128 × 20 × 20
SPP and B2/B3/B4 feed the neck.
SPP
B3
B4
Top-down and bottom-up neck
Upsample ×2
128 × 40 × 40
Concat with B4
224 × 40 × 40
RepNCSPELAN (N4)
96 × 40 × 40; part=96
Upsample ×2
96 × 80 × 80
Concat with B3
160 × 80 × 80
RepNCSPELAN (P3)
64 × 80 × 80; part=64
AConv
48 × 40 × 40
Concat with N4
144 × 40 × 40
RepNCSPELAN (P4)
96 × 40 × 40; part=96
AConv
64 × 20 × 20
Concat with SPP
192 × 20 × 20
RepNCSPELAN (P5)
128 × 20 × 20; part=128
SPP from backbone
B4
B3
N4
SPP
N4
Nearest upsampling. Concat uses channels.
RepNCSP repeat count n = 3.
DDetect head
P3 from neck; stride 8
Input
64 × 80 × 80
Conv 3×3
64 ch
Conv 3×3
64 ch; g=4
Conv2d 1×1
64 ch; g=4
Conv 3×3
80 ch
Conv 3×3
80 ch
Conv2d 1×1
80 ch
Concat box distributions and class logits
144 × 80 × 80; 6,400 locations
P4 from neck; stride 16
Input
96 × 40 × 40
Conv 3×3
64 ch
Conv 3×3
64 ch; g=4
Conv2d 1×1
64 ch; g=4
Conv 3×3
80 ch
Conv 3×3
80 ch
Conv2d 1×1
80 ch
Concat box distributions and class logits
144 × 40 × 40; 1,600 locations
P5 from neck; stride 32
Input
128 × 20 × 20
Conv 3×3
64 ch
Conv 3×3
64 ch; g=4
Conv2d 1×1
64 ch; g=4
Conv 3×3
80 ch
Conv 3×3
80 ch
Conv2d 1×1
80 ch
Concat box distributions and class logits
144 × 20 × 20; 400 locations
Raw location count: 8,400. Decode gives 1 × 84 × 8,400.
Final 1×1 layers have bias, no normalization or activation.
Confidence filtering and NMS run after neural-network decoding.
Conv
Conv2d
k, stride, output channels, groups given by occurrence
BatchNorm2d
eps=0.001; momentum=0.03
SiLU
x × sigmoid(x)
Convolution has no bias. Padding is k//2 for shown odd kernels.
Unmarked strides and group counts are 1.
AConv
AvgPool2d
k=2, s=1, p=0
Conv 3×3
s=2, p=1; output channels from occurrence
Square sizes: 160 to 159 to 80; 80 to 79 to 40;
40 to 39 to 20.
Input channels: 32 / 64 / 96 / 64 / 96
Output channels: 64 / 96 / 128 / 48 / 64
RepConvN
Input
16 / 24 / 32 channels
Conv2d 3×3
p=1; no bias
Conv2d 1×1
p=0; no bias
BatchNorm2d
eps=0.001
BatchNorm2d
eps=0.001
+
SiLU
16 / 24 / 32 channels
The optional identity BatchNorm branch is disabled.
RepNBottleneck
Input
16 / 24 / 32 channels
RepConvN
16 / 24 / 32 channels
Conv 3×3
16 / 24 / 32 channels
+
identity
Equal input/output widths; the residual is enabled.
RepNCSPELAN
Part widths by occurrence: 64 / 96 / 128
Conv 1×1
64 / 96 / 128 channels
Split
32 / 48 / 64 channels per half
RepNCSP (n=3)
32 / 48 / 64 channels
Conv 3×3
32 / 48 / 64 channels
RepNCSP (n=3)
32 / 48 / 64 channels
Conv 3×3
32 / 48 / 64 channels
Concat four inputs
128 / 192 / 256 channels
Conv 1×1
Output width of its stage
RepNCSP
Input
32 / 48 / 64 channels
Conv 1×1
16 / 24 / 32 channels
Conv 1×1
16 / 24 / 32 channels
RepNBottleneck
16 / 24 / 32 channels
RepNBottleneck
16 / 24 / 32 channels
RepNBottleneck
16 / 24 / 32 channels
Concat
32 / 48 / 64 channels
Conv 1×1
32 / 48 / 64 channels
SPPELAN
20 × 20 spatial grid throughout.
Conv 1×1
64 channels
MaxPool2d
k=5, s=1, p=2; 64 ch
MaxPool2d
k=5, s=1, p=2; 64 ch
MaxPool2d
k=5, s=1, p=2; 64 ch
Concat four taps
256 channels
Conv 1×1
128 channels
Decode
Flatten spatial axes and join scales
64 × 8,400 box logits; 80 × 8,400 class logits
Reshape; softmax over 16 bins
4 × 16 × 8,400
Sigmoid
80 × 8,400
Weighted sum over bins 0...15
4 × 8,400 distances
Grid centers minus/plus distances
xyxy corners; scale by feature stride
Concat boxes and class scores
1 × 84 × 8,400
Confidence filtering and NMS follow decode.
ELAN
Conv 1×1
32 channels
Split
16 channels per half
Conv 3×3
16 channels
Conv 3×3
16 channels
Concat four inputs
64 channels
Conv 1×1
32 channels
RepNCSPELAN part widths: B3=64, B4=96, B5=128, N4=96, P3=64, P4=96, P5=128
Eval graph only. Optional PGI or dual-assignment training branches are not executed. Shape checks use random weights.
Source: models/yolo9/nn.py; models/yolo9/nn.py. Revision a4d0ecc9e17f.
libreyolo.com