PP-YOLOE-S
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
PP-YOLOE-S
Detection; 640 × 640 RGB; 80 classes; batch 1; unfused eval. Head order is stride 32,16,8.
LibreYOLO
PP-YOLOE-S
Detection; 640 × 640 RGB; 80 classes; batch 1; unfused eval. Head order is stride 32,16,8.
CSPResNet backbone
Input
3 × 640 × 640
ConvBNAct 3×3, s=2
16 × 320 × 320
ConvBNAct 3×3
16 × 320 × 320
ConvBNAct 3×3
32 × 320 × 320
CSPResStage B2
64 ×160 ×160; n=1; mid=48
CSPResStage B3
128 ×80 ×80; n=2; mid=96
CSPResStage B4
256 ×40 ×40; n=2; mid=192
CSPResStage B5
512 ×20 ×20; n=1; mid=384
Every backbone stage downsamples by 2 and uses EffectiveSE.
PP-YOLOE CSP-PAN
CSPStage with SPP
384 ×20 ×20; blocks=1
ConvBNAct 1×1
192 ×20 ×20
Nearest upsample ×2
192 ×40 ×40
Concat with B4
448 ×40 ×40
CSPStage
192 ×40 ×40; blocks=1
ConvBNAct 1×1
96 ×40 ×40
Nearest upsample ×2
96 ×80 ×80
Concat with B3
224 ×80 ×80
CSPStage
96 ×80 ×80; blocks=1
ConvBNAct 3×3, s=2
96 ×40 ×40
Concat with N4
288 ×40 ×40
CSPStage
192 ×40 ×40; blocks=1
ConvBNAct 3×3, s=2
192 ×20 ×20
Concat with N5
576 ×20 ×20
CSPStage
384 ×20 ×20; blocks=1
B5
B4
B3
N4
N5
Head receives P5, P4, P3, in that order.
Neck basic blocks do not add residuals.
Efficient Task-aligned head
One scale feature
P5/P4/P3 channels: 384/192/96
AdaptiveAvgPool2d
1 ×1; shared mean for both stems
ESEAttn class stem
Same channel width as input
ESEAttn regression stem
Same channel width as input
+
Conv2d 3×3, p=1
80 logits; bias=True
Conv2d 3×3, p=1
68 logits =4 × 17 bins; bias=True
Sigmoid
1 ×8,400 ×80 scores
DFL expectation over 17 bins
Bins 0...16 ; 4 distances/location
Grid point minus/plus distances
Multiply per-level stride;1 ×8,400 ×4
Independent heads per level; no objectness term.
Flatten order: 400 locations, then 1,600, then 6,400.
NMS is external to this decoded eval graph.
ConvBNAct
Conv2d
Kernel/stride/padding from occurrence; no bias
BatchNorm2d
eps=.00001; momentum=.1
SiLU
x ×sigmoid(x)
Stride 1 unless marked. Conv1×1 p=0; Conv3×3 p=1.
RepVGGBlock
Input
Width unchanged in all these blocks
Conv2d3×3
p=1; no bias
BatchNorm2d
eps=.00001
Conv2d1×1
p=0; no bias
BatchNorm2d
eps=.00001
+
SiLU
No identity BN branch, SE or learnable alpha.
BasicBlock
Input
Backbone widths 24/48/96/192; neck widths from table
ConvBNAct 3×3
Same input/output width
RepVGGBlock
Same width; no stride change
+
Backbone adds identity. Neck returns RepVGG output directly.
CSPResStage
ConvBNAct 3×3, s=2
Mid channels 48/96/192/384
ConvBNAct 1×1
24/48/96/192 channels
ConvBNAct 1×1
24/48/96/192 channels
BasicBlock repeated stage n
Residual enabled
Concat two branches
Restore mid channels
EffectiveSE
Same mid channels
ConvBNAct 1×1
Output channels from backbone stage
CSPStage
Input
Neck-stage channels from main graph
ConvBNAct 1×1
48/96/192 channels
ConvBNAct 1×1
48/96/192 channels
BasicBlock ×1
No residual; 48/96/192 channels
Concat two branches
Output channels of neck stage
ConvBNAct 1×1
Same output width
CSPStageSPP
Input
Neck-stage channels from main graph
ConvBNAct 1×1
192 channels
ConvBNAct 1×1
192 channels
BasicBlock ×1
No residual; 192 channels
SPP (parallel pools)
192 channels
Concat two branches
Output channels of neck stage
ConvBNAct 1×1
Same output width
EffectiveSE
Input
Backbone mid channels
Spatial mean
1 ×1 per channel
Conv2d 1×1
Same channel width; bias=True
Hardsigmoid
clamp(x+3,0,6) /6
Multiply original input by gate
Spatially broadcast channel weights
ESEAttn
Shared spatial mean from head
1 ×1; feature channel width
Conv2d 1×1
Same channel width; bias=True
Sigmoid
Channel gate
Multiply original feature by gate
Original feature is the second input
ConvBNAct 1×1
Same feature channel width
Feature
SPP
Input
192 channels;20 ×20
MaxPool5×5
s=1; p=2
MaxPool9×9
s=1; p=4
MaxPool13×13
s=1; p=6
Concat input and three pooled tensors
768 channels
ConvBNAct 1×1
192 channels
DFL and output contract
Reshape box logits
4 sides × 17 bins ×locations
Softmax over 17 bins
Per side and location
Multiply by bins 0...16; sum
Four l/t/r/b distances
Center point ±distances; ×stride
Offset 0.5; strides32/16/8
Raw diagnostics also include class logits[1,8400,80], distributions[1,8400,68], anchors[8400,4], points[8400,2], counts[400,1600,6400], strides[8400,1].
Backbone block widths: 24/48/96/192. Neck block widths: 48, 96, 192. Raw anchor squares have sizes 160/80/40 pixels.
Source: models/ppyoloe/nn.py. Revision a4d0ecc9e17f.
libreyolo.com