PP-LiteSeg T50

Click a block to read its description, or select it with Tab and Enter.

PP-LiteSeg T50Cityscapes semantic segmentation, 19 classes, 512 × 1024 RGB, native eval. Shapes exclude batch.LibreYOLOPP-LiteSeg T50Cityscapes semantic segmentation, 19 classes, 512 × 1024 RGB, native eval. Shapes exclude batch.STDC backboneRGB + ImageNet normalization3 × 512 × 1024ConvBNReLU 3×33 to 32; s2,p1; 256 × 512ConvBNReLU 3×332 to 64; s2,p1; 128 × 256STDC stage s8, n=2256 × 64 × 128; first stride2F8STDC stage s16, n=2512 × 32 × 64; first stride2F16STDC stage s32, n=21024 × 16 × 32; first stride2F32Every STDC stage begins with one downsampling block.Remaining blocks keep spatial resolution and channel width.50 / 75 indicate source validation scale factors.They are not channel-width multipliers.Projection and context pathsF8256 × 64 × 128ConvBNReLU 3×3256 to 64; s1,p1; P8F16512 × 32 × 64ConvBNReLU 3×3512 to 128; s1,p1; P16F321024 × 16 × 32ConvBNReLU 3×31024 to 128; s1,p1; P32F32 context input1024 × 16 × 32SPPM128 × 16 × 32P8/P16/P32 are independent skips; context reads raw F32.ConvBNReLUConv2dParameters specified at each occurrenceBatchNorm2dReLUUAFM decoder and logitsUAFM 1128 to 128; skip P32; resize ×1Fused feature128 × 16 × 32UAFM 2128 to 64; skip P16; resize ×2Fused feature64 × 32 × 64UAFM 364 to 32; skip P8; resize ×2Fused feature32 × 64 × 128ConvBNReLU 3×332 to 32; s1,p1Dropout 0.0IdentityConv2d 1×132 to 19; bias=FalseBilinear resize ×819 × 512 × 1024Native eval returns only main logits.Three auxiliary heads are training-only and do not execute.SPPM: simple pyramid poolingAdaptive average pool1024 × 1 × 1ConvBNReLU 1×11024 to 128Bilinear resize128 × 16 × 32Adaptive average pool1024 × 2 × 2ConvBNReLU 1×11024 to 128Bilinear resize128 × 16 × 32Adaptive average pool1024 × 4 × 4ConvBNReLU 1×11024 to 128Bilinear resize128 × 16 × 32++ConvBNReLU 3×3128 to 128; s1,p1; context outputEach pooling branch independently reads F32. Three 128-channel results are added, not concatenated.STDC 256, first block, stride2Input64 channelsConvBNReLU 1×164 to 128; stride1AvgPool 3×3stride2, padding1a128 chDepthwise Conv + BNk3,s2,p1; 128 groups; no ReLUConvBNReLU 3×3128 to 64; s1,p1b64 chConvBNReLU 3×364 to 32; s1,p1c32 chConvBNReLU 3×332 to 32; s1,p1d32 chabcdConcat a,b,c,d128+64+32+32 = 256 channelsNamed a,b,c,d connections keep all four inputs distinct.Each later convolution consumes only the previous result.STDC 256, repeated block, stride1Input256 channelsConvBNReLU 1×1256 to 128; stride1Identity a128 channelsa128 chIdentity trunk128 channelsConvBNReLU 3×3128 to 64; s1,p1b64 chConvBNReLU 3×364 to 32; s1,p1c32 chConvBNReLU 3×332 to 32; s1,p1d32 chabcdConcat a,b,c,d128+64+32+32 = 256 channelsNamed a,b,c,d connections keep all four inputs distinct.Each later convolution consumes only the previous result.STDC 512, first block, stride2Input256 channelsConvBNReLU 1×1256 to 256; stride1AvgPool 3×3stride2, padding1a256 chDepthwise Conv + BNk3,s2,p1; 256 groups; no ReLUConvBNReLU 3×3256 to 128; s1,p1b128 chConvBNReLU 3×3128 to 64; s1,p1c64 chConvBNReLU 3×364 to 64; s1,p1d64 chabcdConcat a,b,c,d256+128+64+64 = 512 channelsNamed a,b,c,d connections keep all four inputs distinct.Each later convolution consumes only the previous result.STDC 512, repeated block, stride1Input512 channelsConvBNReLU 1×1512 to 256; stride1Identity a256 channelsa256 chIdentity trunk256 channelsConvBNReLU 3×3256 to 128; s1,p1b128 chConvBNReLU 3×3128 to 64; s1,p1c64 chConvBNReLU 3×364 to 64; s1,p1d64 chabcdConcat a,b,c,d256+128+64+64 = 512 channelsNamed a,b,c,d connections keep all four inputs distinct.Each later convolution consumes only the previous result.STDC 1024, first block, stride2Input512 channelsConvBNReLU 1×1512 to 512; stride1AvgPool 3×3stride2, padding1a512 chDepthwise Conv + BNk3,s2,p1; 512 groups; no ReLUConvBNReLU 3×3512 to 256; s1,p1b256 chConvBNReLU 3×3256 to 128; s1,p1c128 chConvBNReLU 3×3128 to 128; s1,p1d128 chabcdConcat a,b,c,d512+256+128+128 = 1024 channelsNamed a,b,c,d connections keep all four inputs distinct.Each later convolution consumes only the previous result.STDC 1024, repeated block, stride1Input1024 channelsConvBNReLU 1×11024 to 512; stride1Identity a512 channelsa512 chIdentity trunk512 channelsConvBNReLU 3×3512 to 256; s1,p1b256 chConvBNReLU 3×3256 to 128; s1,p1c128 chConvBNReLU 3×3128 to 128; s1,p1d128 chabcdConcat a,b,c,d512+256+128+128 = 1024 channelsNamed a,b,c,d connections keep all four inputs distinct.Each later convolution consumes only the previous result.UAFM 1: 128 to 128 channelsDecoder input X128 channelsProjected skip S128 channelsIdentity resize128 × 16 × 32Identity skip projection128 × 16 × 32Channel mean and maxTwo 1-channel spatial mapsChannel mean and maxTwo 1-channel spatial mapsConcat mean(X),max(X),mean(S),max(S)4 spatial channelsConvBNReLU 3×34 to 2; s1,p1Conv2d 3×3 + BatchNorm2 to 1; s1,p1; no activationSigmoidSpatial attention A: 1 channelMultiply X × A128 channelsMultiply S × (1-A)128 channelsX and S denote the resized decoder and skip tensors above.+ConvBNReLU 3×3128 to 128; s1,p1; output 128 × 16 × 32Projection skips are identities for these registered recipes.Attention uses spatial channel statistics, not token attention.UAFM 2: 128 to 64 channelsDecoder input X128 channelsProjected skip S128 channelsBilinear resize ×2128 × 32 × 64Identity skip projection128 × 32 × 64Channel mean and maxTwo 1-channel spatial mapsChannel mean and maxTwo 1-channel spatial mapsConcat mean(X),max(X),mean(S),max(S)4 spatial channelsConvBNReLU 3×34 to 2; s1,p1Conv2d 3×3 + BatchNorm2 to 1; s1,p1; no activationSigmoidSpatial attention A: 1 channelMultiply X × A128 channelsMultiply S × (1-A)128 channelsX and S denote the resized decoder and skip tensors above.+ConvBNReLU 3×3128 to 64; s1,p1; output 64 × 32 × 64Projection skips are identities for these registered recipes.Attention uses spatial channel statistics, not token attention.UAFM 3: 64 to 32 channelsDecoder input X64 channelsProjected skip S64 channelsBilinear resize ×264 × 64 × 128Identity skip projection64 × 64 × 128Channel mean and maxTwo 1-channel spatial mapsChannel mean and maxTwo 1-channel spatial mapsConcat mean(X),max(X),mean(S),max(S)4 spatial channelsConvBNReLU 3×34 to 2; s1,p1Conv2d 3×3 + BatchNorm2 to 1; s1,p1; no activationSigmoidSpatial attention A: 1 channelMultiply X × A64 channelsMultiply S × (1-A)64 channelsX and S denote the resized decoder and skip tensors above.+ConvBNReLU 3×364 to 32; s1,p1; output 32 × 64 × 128Projection skips are identities for these registered recipes.Attention uses spatial channel statistics, not token attention.ConvBNReLU = Conv2d, BatchNorm2d, ReLU in sequence. Numeric channels/kernels at each occurrence; bias=False.Source: libreyolo/models/ppliteseg/nn.py and model.py. Revision a4d0ecc9e17f.libreyolo.com