PicoDet-S
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
PicoDet-S
320 × 320 RGB; 80 classes; batch 1. Unfused PyTorch eval; four feature scales.
LibreYOLO
PicoDet-S
320 × 320 RGB; 80 classes; batch 1. Unfused PyTorch eval; four feature scales.
ESNet backbone
Input
3 × 320 × 320
ConvBNAct 3×3, s=2
24 × 160 × 160
MaxPool 3×3, s=2, p=1
24 × 80 × 80
Block 0: ESBlockDS; mid=88
24 input; 96 × 40 × 40
Block 1: ESBlock; mid=48
96 input; 96 × 40 × 40
Block 2: ESBlock; mid=48
96 input; 96 × 40 × 40; output C3
Block 3: ESBlockDS; mid=96
96 input; 192 × 20 × 20
Block 4: ESBlock; mid=120
192 input; 192 × 20 × 20
Block 5: ESBlock; mid=96
192 input; 192 × 20 × 20
Block 6: ESBlock; mid=120
192 input; 192 × 20 × 20
Block 7: ESBlock; mid=96
192 input; 192 × 20 × 20
Block 8: ESBlock; mid=96
192 input; 192 × 20 × 20
Block 9: ESBlock; mid=96
192 input; 192 × 20 × 20; output C4
Block 10: ESBlockDS; mid=192
192 input; 384 × 10 × 10
Block 11: ESBlock; mid=192
384 input; 384 × 10 × 10
Block 12: ESBlock; mid=192
384 input; 384 × 10 × 10; output C5
Downsample blocks: 0,3,10. Outputs: 2,9,12.
CSP-PAN
C3: ConvBNAct 1×1
96 input; 96 × 40 × 40
C4: ConvBNAct 1×1
192 input; 96 × 20 × 20
C5: ConvBNAct 1×1
384 input; 96 × 10 × 10
Nearest resize ×2
96 × 20 × 20
Concat with T4
192 × 20 × 20
CSPLayer, n=1
96 × 20 × 20
Nearest resize ×2
96 × 40 × 40
Concat with T3
192 × 40 × 40
CSPLayer, n=1
96 × 40 × 40
DepthwiseSeparable 5×5, s=2
96 × 20 × 20
Concat with N4
192 × 20 × 20
CSPLayer, n=1
96 × 20 × 20
DepthwiseSeparable 5×5, s=2
96 × 10 × 10
Concat with T5
192 × 10 × 10
CSPLayer, n=1
96 × 10 × 10
T5
T4
T3
N4
T5
T5: DepthwiseSeparable 5×5, s=2
96 × 5 × 5
P5: DepthwiseSeparable 5×5, s=2
96 × 5 × 5
+
P6 output
96 × 5 × 5
All four CSP neck blocks disable residual addition.
PicoHead and decoding
One scale feature
96 channels; execute independently at P3/P4/P5/P6
DepthwiseSeparable 5×5
96 channels; stack 2 total layers
DepthwiseSeparable 5×5
96 channels; stack 2 total layers
Conv2d 1×1
112 output channels; bias=True
Split channels
80 class logits; 32 box-distribution logits
Sigmoid class logits
80 probabilities/location
Softmax over 8 bins per side
Weighted expectation of bins 0...7
Multiply distances by stride
8,16,32,64 pixels
Grid center minus/plus l/t/r/b
Grid offset=0.5; xyxy boxes
Threshold and class-aware NMS
Decode and NMS are outside native raw-head forward
No separate regression tower or objectness branch.
ConvBNAct
Conv2d
Bias=False; p=k//2; k/stride/groups from occurrence
BatchNorm2d
eps=.00001; momentum=.1
Hardswish or identity
Identity only on marked ES depthwise branches
ConvBNAct depthwise k×k
groups=input channels; k=5 in neck/head
ConvBNAct pointwise 1×1
groups=1; output width from occurrence
The lower pair defines DepthwiseSeparableConv.
ESBlock
Input
Cin channels
Split into two channel halves
A channels each; table resolves A
ConvBNAct 1×1
A input; B output; Hardswish
ConvBNAct depthwise 3×3
B channels; groups=B; no activation
Concat PW and DW outputs
G channels
SELayer
G input; R reduced width
ConvBNAct 1×1
G input; O output; Hardswish
Concat untouched half and new branch
Cout channels
Reshape
1 ×2 ×O ×height ×width
Transpose channel-group axes 1 and 2
Make contiguous
Reshape
1 ×Cout ×height ×width
ESBlockDS
Input
Cin channels; no channel split before branches
Depthwise 3×3, s=2
Cin channels; BN; no activation
ConvBNAct 1×1
Cin input; O output; Hardswish
ConvBNAct 1×1
Cin input; B output; Hardswish
Depthwise 3×3, s=2
B channels; BN; no activation
SELayer
G=B input; R reduced width
ConvBNAct 1×1
B input; O output; Hardswish
Concat two downsampled branches
Cout channels
ConvBNAct depthwise 3×3
Cout channels; s=1; BN + Hardswish
ConvBNAct pointwise 1×1
Cout channels; BN + Hardswish
This downsample block does not apply channel shuffle.
SELayer
Input
G channels
AdaptiveAvgPool2d
1 ×1 spatial output
Conv2d 1×1
G input; R output; bias=True
ReLU
Conv2d 1×1
R input; G output; bias=True
HSigmoid (custom)
clamp((x+3)/6,0,6); range[0,6]
Multiply original feature by gate
Broadcast across spatial positions
This source gate is not the standard [0,1] hardsigmoid.
CSPLayer
Input
192 channels
ConvBNAct 1×1
48 channels
ConvBNAct 1×1
48 channels
DarknetBottleneck, n=1
No residual; hidden 48
Concat
96 channels
ConvBNAct 1×1
96 output channels
DarknetBottleneck
Input
48 channels
ConvBNAct 1×1
Width unchanged
DepthwiseSeparable 5×5
Width unchanged; BN and Hardswish after both convs
No residual addition in the configured PicoDet CSP-PAN.
Per-block width and resolution values
Block
Cin
Cout
Mid M
A
B
G
R
O
Grid
Type
0
24
96
88
24
44
44
11
48
40
DS
1
96
96
48
48
24
48
12
48
40
ES
2
96
96
48
48
24
48
12
48
40
ES
3
96
192
96
96
48
48
12
96
20
DS
4
192
192
120
96
60
120
30
96
20
ES
5
192
192
96
96
48
96
24
96
20
ES
6
192
192
120
96
60
120
30
96
20
ES
7
192
192
96
96
48
96
24
96
20
ES
8
192
192
96
96
48
96
24
96
20
ES
9
192
192
96
96
48
96
24
96
20
ES
10
192
384
192
192
96
96
24
192
10
DS
11
384
384
192
192
96
192
48
192
10
ES
12
384
384
192
192
96
192
48
192
10
ES
Source: models/picodet/nn.py; postprocess/picodet.py. Revision a4d0ecc9e17f.
libreyolo.com