PP-LiteSeg T50
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
PP-LiteSeg T50
Cityscapes semantic segmentation, 19 classes, 512 × 1024 RGB, native eval. Shapes exclude batch.
LibreYOLO
PP-LiteSeg T50
Cityscapes semantic segmentation, 19 classes, 512 × 1024 RGB, native eval. Shapes exclude batch.
STDC backbone
RGB + ImageNet normalization
3 × 512 × 1024
ConvBNReLU 3×3
3 to 32; s2,p1; 256 × 512
ConvBNReLU 3×3
32 to 64; s2,p1; 128 × 256
STDC stage s8, n=2
256 × 64 × 128; first stride2
F8
STDC stage s16, n=2
512 × 32 × 64; first stride2
F16
STDC stage s32, n=2
1024 × 16 × 32; first stride2
F32
Every STDC stage begins with one downsampling block.
Remaining blocks keep spatial resolution and channel width.
50 / 75 indicate source validation scale factors.
They are not channel-width multipliers.
Projection and context paths
F8
256 × 64 × 128
ConvBNReLU 3×3
256 to 64; s1,p1; P8
F16
512 × 32 × 64
ConvBNReLU 3×3
512 to 128; s1,p1; P16
F32
1024 × 16 × 32
ConvBNReLU 3×3
1024 to 128; s1,p1; P32
F32 context input
1024 × 16 × 32
SPPM
128 × 16 × 32
P8/P16/P32 are independent skips; context reads raw F32.
ConvBNReLU
Conv2d
Parameters specified at each occurrence
BatchNorm2d
ReLU
UAFM decoder and logits
UAFM 1
128 to 128; skip P32; resize ×1
Fused feature
128 × 16 × 32
UAFM 2
128 to 64; skip P16; resize ×2
Fused feature
64 × 32 × 64
UAFM 3
64 to 32; skip P8; resize ×2
Fused feature
32 × 64 × 128
ConvBNReLU 3×3
32 to 32; s1,p1
Dropout 0.0
Identity
Conv2d 1×1
32 to 19; bias=False
Bilinear resize ×8
19 × 512 × 1024
Native eval returns only main logits.
Three auxiliary heads are training-only and do not execute.
SPPM: simple pyramid pooling
Adaptive average pool
1024 × 1 × 1
ConvBNReLU 1×1
1024 to 128
Bilinear resize
128 × 16 × 32
Adaptive average pool
1024 × 2 × 2
ConvBNReLU 1×1
1024 to 128
Bilinear resize
128 × 16 × 32
Adaptive average pool
1024 × 4 × 4
ConvBNReLU 1×1
1024 to 128
Bilinear resize
128 × 16 × 32
+
+
ConvBNReLU 3×3
128 to 128; s1,p1; context output
Each pooling branch independently reads F32. Three 128-channel results are added, not concatenated.
STDC 256, first block, stride2
Input
64 channels
ConvBNReLU 1×1
64 to 128; stride1
AvgPool 3×3
stride2, padding1
a
128 ch
Depthwise Conv + BN
k3,s2,p1; 128 groups; no ReLU
ConvBNReLU 3×3
128 to 64; s1,p1
b
64 ch
ConvBNReLU 3×3
64 to 32; s1,p1
c
32 ch
ConvBNReLU 3×3
32 to 32; s1,p1
d
32 ch
a
b
c
d
Concat a,b,c,d
128+64+32+32 = 256 channels
Named a,b,c,d connections keep all four inputs distinct.
Each later convolution consumes only the previous result.
STDC 256, repeated block, stride1
Input
256 channels
ConvBNReLU 1×1
256 to 128; stride1
Identity a
128 channels
a
128 ch
Identity trunk
128 channels
ConvBNReLU 3×3
128 to 64; s1,p1
b
64 ch
ConvBNReLU 3×3
64 to 32; s1,p1
c
32 ch
ConvBNReLU 3×3
32 to 32; s1,p1
d
32 ch
a
b
c
d
Concat a,b,c,d
128+64+32+32 = 256 channels
Named a,b,c,d connections keep all four inputs distinct.
Each later convolution consumes only the previous result.
STDC 512, first block, stride2
Input
256 channels
ConvBNReLU 1×1
256 to 256; stride1
AvgPool 3×3
stride2, padding1
a
256 ch
Depthwise Conv + BN
k3,s2,p1; 256 groups; no ReLU
ConvBNReLU 3×3
256 to 128; s1,p1
b
128 ch
ConvBNReLU 3×3
128 to 64; s1,p1
c
64 ch
ConvBNReLU 3×3
64 to 64; s1,p1
d
64 ch
a
b
c
d
Concat a,b,c,d
256+128+64+64 = 512 channels
Named a,b,c,d connections keep all four inputs distinct.
Each later convolution consumes only the previous result.
STDC 512, repeated block, stride1
Input
512 channels
ConvBNReLU 1×1
512 to 256; stride1
Identity a
256 channels
a
256 ch
Identity trunk
256 channels
ConvBNReLU 3×3
256 to 128; s1,p1
b
128 ch
ConvBNReLU 3×3
128 to 64; s1,p1
c
64 ch
ConvBNReLU 3×3
64 to 64; s1,p1
d
64 ch
a
b
c
d
Concat a,b,c,d
256+128+64+64 = 512 channels
Named a,b,c,d connections keep all four inputs distinct.
Each later convolution consumes only the previous result.
STDC 1024, first block, stride2
Input
512 channels
ConvBNReLU 1×1
512 to 512; stride1
AvgPool 3×3
stride2, padding1
a
512 ch
Depthwise Conv + BN
k3,s2,p1; 512 groups; no ReLU
ConvBNReLU 3×3
512 to 256; s1,p1
b
256 ch
ConvBNReLU 3×3
256 to 128; s1,p1
c
128 ch
ConvBNReLU 3×3
128 to 128; s1,p1
d
128 ch
a
b
c
d
Concat a,b,c,d
512+256+128+128 = 1024 channels
Named a,b,c,d connections keep all four inputs distinct.
Each later convolution consumes only the previous result.
STDC 1024, repeated block, stride1
Input
1024 channels
ConvBNReLU 1×1
1024 to 512; stride1
Identity a
512 channels
a
512 ch
Identity trunk
512 channels
ConvBNReLU 3×3
512 to 256; s1,p1
b
256 ch
ConvBNReLU 3×3
256 to 128; s1,p1
c
128 ch
ConvBNReLU 3×3
128 to 128; s1,p1
d
128 ch
a
b
c
d
Concat a,b,c,d
512+256+128+128 = 1024 channels
Named a,b,c,d connections keep all four inputs distinct.
Each later convolution consumes only the previous result.
UAFM 1: 128 to 128 channels
Decoder input X
128 channels
Projected skip S
128 channels
Identity resize
128 × 16 × 32
Identity skip projection
128 × 16 × 32
Channel mean and max
Two 1-channel spatial maps
Channel mean and max
Two 1-channel spatial maps
Concat mean(X),max(X),mean(S),max(S)
4 spatial channels
ConvBNReLU 3×3
4 to 2; s1,p1
Conv2d 3×3 + BatchNorm
2 to 1; s1,p1; no activation
Sigmoid
Spatial attention A: 1 channel
Multiply X × A
128 channels
Multiply S × (1-A)
128 channels
X and S denote the resized decoder and skip tensors above.
+
ConvBNReLU 3×3
128 to 128; s1,p1; output 128 × 16 × 32
Projection skips are identities for these registered recipes.
Attention uses spatial channel statistics, not token attention.
UAFM 2: 128 to 64 channels
Decoder input X
128 channels
Projected skip S
128 channels
Bilinear resize ×2
128 × 32 × 64
Identity skip projection
128 × 32 × 64
Channel mean and max
Two 1-channel spatial maps
Channel mean and max
Two 1-channel spatial maps
Concat mean(X),max(X),mean(S),max(S)
4 spatial channels
ConvBNReLU 3×3
4 to 2; s1,p1
Conv2d 3×3 + BatchNorm
2 to 1; s1,p1; no activation
Sigmoid
Spatial attention A: 1 channel
Multiply X × A
128 channels
Multiply S × (1-A)
128 channels
X and S denote the resized decoder and skip tensors above.
+
ConvBNReLU 3×3
128 to 64; s1,p1; output 64 × 32 × 64
Projection skips are identities for these registered recipes.
Attention uses spatial channel statistics, not token attention.
UAFM 3: 64 to 32 channels
Decoder input X
64 channels
Projected skip S
64 channels
Bilinear resize ×2
64 × 64 × 128
Identity skip projection
64 × 64 × 128
Channel mean and max
Two 1-channel spatial maps
Channel mean and max
Two 1-channel spatial maps
Concat mean(X),max(X),mean(S),max(S)
4 spatial channels
ConvBNReLU 3×3
4 to 2; s1,p1
Conv2d 3×3 + BatchNorm
2 to 1; s1,p1; no activation
Sigmoid
Spatial attention A: 1 channel
Multiply X × A
64 channels
Multiply S × (1-A)
64 channels
X and S denote the resized decoder and skip tensors above.
+
ConvBNReLU 3×3
64 to 32; s1,p1; output 32 × 64 × 128
Projection skips are identities for these registered recipes.
Attention uses spatial channel statistics, not token attention.
ConvBNReLU = Conv2d, BatchNorm2d, ReLU in sequence. Numeric channels/kernels at each occurrence; bias=False.
Source: libreyolo/models/ppliteseg/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com