PP-YOLOE-S

Click a block to read its description, or select it with Tab and Enter.

PP-YOLOE-SDetection; 640 × 640 RGB; 80 classes; batch 1; unfused eval. Head order is stride 32,16,8.LibreYOLOPP-YOLOE-SDetection; 640 × 640 RGB; 80 classes; batch 1; unfused eval. Head order is stride 32,16,8.CSPResNet backboneInput3 × 640 × 640ConvBNAct 3×3, s=216 × 320 × 320ConvBNAct 3×316 × 320 × 320ConvBNAct 3×332 × 320 × 320CSPResStage B264 ×160 ×160; n=1; mid=48CSPResStage B3128 ×80 ×80; n=2; mid=96CSPResStage B4256 ×40 ×40; n=2; mid=192CSPResStage B5512 ×20 ×20; n=1; mid=384Every backbone stage downsamples by 2 and uses EffectiveSE.PP-YOLOE CSP-PANCSPStage with SPP384 ×20 ×20; blocks=1ConvBNAct 1×1192 ×20 ×20Nearest upsample ×2192 ×40 ×40Concat with B4448 ×40 ×40CSPStage192 ×40 ×40; blocks=1ConvBNAct 1×196 ×40 ×40Nearest upsample ×296 ×80 ×80Concat with B3224 ×80 ×80CSPStage96 ×80 ×80; blocks=1ConvBNAct 3×3, s=296 ×40 ×40Concat with N4288 ×40 ×40CSPStage192 ×40 ×40; blocks=1ConvBNAct 3×3, s=2192 ×20 ×20Concat with N5576 ×20 ×20CSPStage384 ×20 ×20; blocks=1B5B4B3N4N5Head receives P5, P4, P3, in that order.Neck basic blocks do not add residuals.Efficient Task-aligned headOne scale featureP5/P4/P3 channels: 384/192/96AdaptiveAvgPool2d1 ×1; shared mean for both stemsESEAttn class stemSame channel width as inputESEAttn regression stemSame channel width as input+Conv2d 3×3, p=180 logits; bias=TrueConv2d 3×3, p=168 logits =4 × 17 bins; bias=TrueSigmoid1 ×8,400 ×80 scoresDFL expectation over 17 binsBins 0...16 ; 4 distances/locationGrid point minus/plus distancesMultiply per-level stride;1 ×8,400 ×4Independent heads per level; no objectness term.Flatten order: 400 locations, then 1,600, then 6,400.NMS is external to this decoded eval graph.ConvBNActConv2dKernel/stride/padding from occurrence; no biasBatchNorm2deps=.00001; momentum=.1SiLUx ×sigmoid(x)Stride 1 unless marked. Conv1×1 p=0; Conv3×3 p=1.RepVGGBlockInputWidth unchanged in all these blocksConv2d3×3p=1; no biasBatchNorm2deps=.00001Conv2d1×1p=0; no biasBatchNorm2deps=.00001+SiLUNo identity BN branch, SE or learnable alpha.BasicBlockInputBackbone widths 24/48/96/192; neck widths from tableConvBNAct 3×3Same input/output widthRepVGGBlockSame width; no stride change+Backbone adds identity. Neck returns RepVGG output directly.CSPResStageConvBNAct 3×3, s=2Mid channels 48/96/192/384ConvBNAct 1×124/48/96/192 channelsConvBNAct 1×124/48/96/192 channelsBasicBlock repeated stage nResidual enabledConcat two branchesRestore mid channelsEffectiveSESame mid channelsConvBNAct 1×1Output channels from backbone stageCSPStageInputNeck-stage channels from main graphConvBNAct 1×148/96/192 channelsConvBNAct 1×148/96/192 channelsBasicBlock ×1No residual; 48/96/192 channelsConcat two branchesOutput channels of neck stageConvBNAct 1×1Same output widthCSPStageSPPInputNeck-stage channels from main graphConvBNAct 1×1192 channelsConvBNAct 1×1192 channelsBasicBlock ×1No residual; 192 channelsSPP (parallel pools)192 channelsConcat two branchesOutput channels of neck stageConvBNAct 1×1Same output widthEffectiveSEInputBackbone mid channelsSpatial mean1 ×1 per channelConv2d 1×1Same channel width; bias=TrueHardsigmoidclamp(x+3,0,6) /6Multiply original input by gateSpatially broadcast channel weightsESEAttnShared spatial mean from head1 ×1; feature channel widthConv2d 1×1Same channel width; bias=TrueSigmoidChannel gateMultiply original feature by gateOriginal feature is the second inputConvBNAct 1×1Same feature channel widthFeatureSPPInput192 channels;20 ×20MaxPool5×5s=1; p=2MaxPool9×9s=1; p=4MaxPool13×13s=1; p=6Concat input and three pooled tensors768 channelsConvBNAct 1×1192 channelsDFL and output contractReshape box logits4 sides × 17 bins ×locationsSoftmax over 17 binsPer side and locationMultiply by bins 0...16; sumFour l/t/r/b distancesCenter point ±distances; ×strideOffset 0.5; strides32/16/8Raw diagnostics also include class logits[1,8400,80], distributions[1,8400,68], anchors[8400,4], points[8400,2], counts[400,1600,6400], strides[8400,1].Backbone block widths: 24/48/96/192. Neck block widths: 48, 96, 192. Raw anchor squares have sizes 160/80/40 pixels.Source: models/ppyoloe/nn.py. Revision a4d0ecc9e17f.libreyolo.com