YOLOX-S
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
YOLOX-S
Detection; 640 × 640 RGB; 80 classes; batch 1. Unfused PyTorch eval; raw box offsets and sigmoid scores.
LibreYOLO
YOLOX-S
Detection; 640 × 640 RGB; 80 classes; batch 1. Unfused PyTorch eval; raw box offsets and sigmoid scores.
CSPDarknet backbone
Input
3 × 640 × 640
Focus
32 × 320 × 320
BaseConv 3×3, s=2
64 × 160 × 160
CSPLayer, residual enabled
64 × 160 × 160; n=1
BaseConv 3×3, s=2
128 × 80 × 80
CSPLayer (B3)
128 × 80 × 80; n=3
BaseConv 3×3, s=2
256 × 40 × 40
CSPLayer (B4)
256 × 40 × 40; n=3
BaseConv 3×3, s=2
512 × 20 × 20
SPPBottleneck
512 × 20 × 20
CSPLayer (B5), no residual
512 × 20 × 20; n=1
B3, B4 and B5 continue to the neck.
YOLOPAFPN
BaseConv 1×1
256 × 20 × 20
Nearest upsample ×2
256 × 40 × 40
Concat with B4
512 × 40 × 40
CSPLayer (no residual)
256 × 40 × 40; n=1
BaseConv 1×1
128 × 40 × 40
Nearest upsample ×2
128 × 80 × 80
Concat with B3
256 × 80 × 80
CSPLayer (P3)
128 × 80 × 80; n=1
BaseConv 3×3, s=2
128 × 40 × 40
Concat with red
256 × 40 × 40
CSPLayer (P4)
256 × 40 × 40; n=1
BaseConv 3×3, s=2
256 × 20 × 20
Concat with lat
512 × 20 × 20
CSPLayer (P5)
512 × 20 × 20; n=1
B5
B4
B3
red
lat
lat
red
Neck bottlenecks do not add residuals.
Matching B/lat/red labels identify tensor continuations.
YOLOXHead (three independent scales)
Execute this graph separately for P3, P4 and P5.
One scale feature
P3 / P4 / P5 dimensions below
BaseConv 1×1
128 output channels
BaseConv 3×3
128 channels
BaseConv 3×3
128 channels
BaseConv 3×3
128 channels
BaseConv 3×3
128 channels
Conv2d 1×1
80 logits; bias=True
Sigmoid
80 class probabilities
Conv2d 1×1
4 box offsets
Conv2d 1×1
1 objectness logit
Sigmoid
1 probability
Concat box offsets, objectness, class probabilities
85 channels per location
Scale feature channels square grid raw output channels
P3
128
80
85
P4
256
40
85
P5
512
20
85
Convolutions in different scales have independent weights.
No DFL bins. Objectness and class probabilities are separate.
BaseConv
Conv2d
k, stride, groups, channels from occurrence; no bias
BatchNorm2d
eps=0.001; momentum=0.03
SiLU
x × sigmoid(x)
Padding = (k - 1) / 2 for the odd kernels shown.
Focus
Input
3 × 640 × 640
Even row, even col
3 channels
Odd row, even col
3 channels
Even row, odd col
3 channels
Odd row, odd col
3 channels
Concat in TL, BL, TR, BR order
12 × 320 × 320
BaseConv 3×3
32 channels
CSPLayer
Input
Width and repeats from occurrence
BaseConv 1×1
32 / 64 / 128 / 256 hidden channels
BaseConv 1×1
Same hidden width
Bottleneck repeated n times
n is printed on each backbone/neck block
Concat
64 / 128 / 256 / 512 channels
BaseConv 1×1
Output width Q of the occurrence
Bottleneck
Input
Hidden width of parent CSPLayer
BaseConv 1×1
Width unchanged
BaseConv 3×3
Width unchanged
+
Residual only in dark2, dark3 and dark4; otherwise output conv2.
SPPBottleneck
BaseConv 1×1
256 channels
MaxPool2d 5×5
s=1; p=2
MaxPool2d 9×9
s=1; p=4
MaxPool2d 13×13
s=1; p=6
Concat input and three parallel pools
1024 channels
BaseConv 1×1
512 channels
Decode
Flatten and concatenate three scales
1 × 8,400 × 85
Add zero-based grid; multiply stride
x/y offsets have no sigmoid
Exp width/height; multiply stride
Strides: 8, 16, 32
Join cx, cy, width, height
Convert to corner coordinates
Objectness × class scores; filter; NMS
Postprocessing occurs after raw network outputs
No anchor templates or half-cell grid offset.
Hidden CSP widths: 32, 64, 128, 256. Backbone repeats: 1, 3, 3, 1; neck repeats: 1.
Source: models/yolox/nn.py; postprocess/yolox.py. Revision a4d0ecc9e17f.
libreyolo.com