YOLOX-S

Click a block to read its description, or select it with Tab and Enter.

YOLOX-SDetection; 640 × 640 RGB; 80 classes; batch 1. Unfused PyTorch eval; raw box offsets and sigmoid scores.LibreYOLOYOLOX-SDetection; 640 × 640 RGB; 80 classes; batch 1. Unfused PyTorch eval; raw box offsets and sigmoid scores.CSPDarknet backboneInput3 × 640 × 640Focus32 × 320 × 320BaseConv 3×3, s=264 × 160 × 160CSPLayer, residual enabled64 × 160 × 160; n=1BaseConv 3×3, s=2128 × 80 × 80CSPLayer (B3)128 × 80 × 80; n=3BaseConv 3×3, s=2256 × 40 × 40CSPLayer (B4)256 × 40 × 40; n=3BaseConv 3×3, s=2512 × 20 × 20SPPBottleneck512 × 20 × 20CSPLayer (B5), no residual512 × 20 × 20; n=1B3, B4 and B5 continue to the neck.YOLOPAFPNBaseConv 1×1256 × 20 × 20Nearest upsample ×2256 × 40 × 40Concat with B4512 × 40 × 40CSPLayer (no residual)256 × 40 × 40; n=1BaseConv 1×1128 × 40 × 40Nearest upsample ×2128 × 80 × 80Concat with B3256 × 80 × 80CSPLayer (P3)128 × 80 × 80; n=1BaseConv 3×3, s=2128 × 40 × 40Concat with red256 × 40 × 40CSPLayer (P4)256 × 40 × 40; n=1BaseConv 3×3, s=2256 × 20 × 20Concat with lat512 × 20 × 20CSPLayer (P5)512 × 20 × 20; n=1B5B4B3redlatlatredNeck bottlenecks do not add residuals.Matching B/lat/red labels identify tensor continuations.YOLOXHead (three independent scales)Execute this graph separately for P3, P4 and P5.One scale featureP3 / P4 / P5 dimensions belowBaseConv 1×1128 output channelsBaseConv 3×3128 channelsBaseConv 3×3128 channelsBaseConv 3×3128 channelsBaseConv 3×3128 channelsConv2d 1×180 logits; bias=TrueSigmoid80 class probabilitiesConv2d 1×14 box offsetsConv2d 1×11 objectness logitSigmoid1 probabilityConcat box offsets, objectness, class probabilities85 channels per locationScale feature channels square grid raw output channelsP31288085P42564085P55122085Convolutions in different scales have independent weights.No DFL bins. Objectness and class probabilities are separate.BaseConvConv2dk, stride, groups, channels from occurrence; no biasBatchNorm2deps=0.001; momentum=0.03SiLUx × sigmoid(x)Padding = (k - 1) / 2 for the odd kernels shown.FocusInput3 × 640 × 640Even row, even col3 channelsOdd row, even col3 channelsEven row, odd col3 channelsOdd row, odd col3 channelsConcat in TL, BL, TR, BR order12 × 320 × 320BaseConv 3×332 channelsCSPLayerInputWidth and repeats from occurrenceBaseConv 1×132 / 64 / 128 / 256 hidden channelsBaseConv 1×1Same hidden widthBottleneck repeated n timesn is printed on each backbone/neck blockConcat64 / 128 / 256 / 512 channelsBaseConv 1×1Output width Q of the occurrenceBottleneckInputHidden width of parent CSPLayerBaseConv 1×1Width unchangedBaseConv 3×3Width unchanged+Residual only in dark2, dark3 and dark4; otherwise output conv2.SPPBottleneckBaseConv 1×1256 channelsMaxPool2d 5×5s=1; p=2MaxPool2d 9×9s=1; p=4MaxPool2d 13×13s=1; p=6Concat input and three parallel pools1024 channelsBaseConv 1×1512 channelsDecodeFlatten and concatenate three scales1 × 8,400 × 85Add zero-based grid; multiply stridex/y offsets have no sigmoidExp width/height; multiply strideStrides: 8, 16, 32Join cx, cy, width, heightConvert to corner coordinatesObjectness × class scores; filter; NMSPostprocessing occurs after raw network outputsNo anchor templates or half-cell grid offset.Hidden CSP widths: 32, 64, 128, 256. Backbone repeats: 1, 3, 3, 1; neck repeats: 1.Source: models/yolox/nn.py; postprocess/yolox.py. Revision a4d0ecc9e17f.libreyolo.com