FeyNobg L
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
FeyNobg L
Alpha-matte logits,normalized RGB input3 × 1024 × 1024,native eval. Shapes exclude batch.
LibreYOLO
FeyNobg L
Alpha-matte logits,normalized RGB input3 × 1024 × 1024,native eval. Shapes exclude batch.
Dual-scale shared encoder
full image
3 × 1024 × 1024
Shared Swin-v1 encoder
E=192;window12
half image
3 × 512 × 512
Shared Swin-v1 encoder
E=192;window12
Half image uses bilinear resize,align_corners=True.
Full stage1
192 × 256²
Resize half stage1
192 × 128² to256²
Concat two scales: X1
384 × 256²
Full stage2
384 × 128²
Resize half stage2
384 × 64² to128²
Concat two scales: X2
768 × 128²
Full stage3
768 × 64²
Resize half stage3
768 × 32² to64²
Concat two scales: X3
1536 × 64²
Full stage4
1536 × 32²
Resize half stage4
1536 × 16² to32²
Concat two scales: X4
3072 × 32²
Resize X1,X2,X3 to32²; concat withX4
5760 × 32 × 32
BasicDecBlk squeeze
5760 to3072;32²
Bilateral-reference decoder
Concat current feature and IPT5
3072 + 384 = 3456 channels at32²
BasicDecBlk
3456 to1536;32²
Gradient-reference attention
1536 channels multiplied by spatial gate
Bilinear resize×2
1536 × 64²
+
Conv1×1(X3)
1536 to1536
Concat current feature and IPT4
1536 + 384 = 1920 channels at64²
BasicDecBlk
1920 to768;64²
Gradient-reference attention
768 channels multiplied by spatial gate
Bilinear resize×2
768 × 128²
+
Conv1×1(X2)
768 to768
Concat current feature and IPT3
768 + 192 = 960 channels at128²
BasicDecBlk
960 to384;128²
Gradient-reference attention
384 channels multiplied by spatial gate
Bilinear resize×2
384 × 256²
+
Conv1×1(X1)
384 to384
Concat current feature and IPT2
384 + 96 = 480 channels at256²
BasicDecBlk
480 to192;256²
Bilinear resize to1024²
192 channels
Concat full-resolution feature and IPT1
192 + 48 = 240 channels
Conv1×1 to one logit
240 to1;1 × 1024 × 1024
Sigmoid to alpha belongs to shared matte postprocessing.
Image-patch reference paths
Each IPT path consumes the original normalized1024 RGB image.
Tile packing for IPT5
32×32 contiguous tiles;3072×32²
Conv3×3 thenConv3×3
3072 to64 to384;boths1,p1
IPT5
384 × 32²
Tile packing for IPT4
16×16 contiguous tiles;768×64²
Conv3×3 thenConv3×3
768 to64 to384;boths1,p1
IPT4
384 × 64²
Tile packing for IPT3
8×8 contiguous tiles;192×128²
Conv3×3 thenConv3×3
192 to64 to192;boths1,p1
IPT3
192 × 128²
Tile packing for IPT2
4×4 contiguous tiles;48×256²
Conv3×3 thenConv3×3
48 to64 to96;boths1,p1
IPT2
96 × 256²
Tile packing for IPT1
1×1 contiguous tiles;3×1024²
Conv3×3 thenConv3×3
3 to64 to48;boths1,p1
IPT1
48 × 1024²
Tile packing concatenates whole spatial tiles as channels;
it is not the pixel-interleaving PixelUnshuffle operation.
Shared Swin-v1 backbone
Conv patch4,stride4 + LayerNorm
3 to192; fullgrid256² / halfgrid128²
Swin block stage1,n=2
192 channels;heads6;MLPwidth768
Output-stage LayerNorm
Full/half grids 256² / 128²
PatchMerging
Concat2×2 gives768;LN;Linear to384
Swin block stage2,n=2
384 channels;heads12;MLPwidth1536
Output-stage LayerNorm
Full/half grids 128² / 64²
PatchMerging
Concat2×2 gives1536;LN;Linear to768
Swin block stage3,n=24
768 channels;heads24;MLPwidth3072
Output-stage LayerNorm
Full/half grids 64² / 32²
PatchMerging
Concat2×2 gives3072;LN;Linear to1536
Swin block stage4,n=2
1536 channels;heads48;MLPwidth6144
Output-stage LayerNorm
Full/half grids 32² / 16²
Stage normalization is an output tap; merging uses pre-tap tokens.
Swin block and attention
Input tokens
Stage widths [192, 384, 768, 1536]
LayerNorm
epsilon1e-5
Pad spatial grid to window multiple
Window12; alternating shift0/6
Partition windows
144 tokens perwindow
Linear QKV; split heads
QKV widths [576, 1152, 2304, 4608]
QK transpose /sqrt32 +relative bias
Learned relative bias; shifted mask0/-100
Softmax thenweights × V
Each head has32 channels
Concat heads; output Linear
Restore stage channel width
Reverse windows,undo shift,crop
Restore unpadded stage grid
+
LN; Linear to4C; GELU; Linear toC
Numeric widths [768, 1536, 3072, 6144]
+
PatchMerging reads2×2 spatial neighbors, thenLN and4C-to2C Linear.
BasicDecBlk
Conv3×3
Occurrence Cin to64;s1,p1,biasTrue
BatchNorm2d + ReLU
64 channels
ASPPDeformable
64 to64
Conv3×3
64 tooccurrence Cout;s1,p1,biasTrue
BatchNorm2d
Cout channels;no finalactivation
Gradient-reference attention
Conv3×3;BN;ReLU
Inputchannels[1536, 768, 384] to16;s1,p1
Conv1×1 +Sigmoid
16to1;spatial gate
×
Feature
Attention executes in eval; gradient-prediction supervision does not.
ASPPDeformable: five parallel branches
DeformableConv1×1
64to256;padding0
BatchNorm2d thenReLU
256 channels
DeformableConv1×1
64to256;padding0
BatchNorm2d thenReLU
256 channels
DeformableConv3×3
64to256;padding1
BatchNorm2d thenReLU
256 channels
DeformableConv7×7
64to256;padding3
BatchNorm2d thenReLU
256 channels
Global avgpool;Conv1×1;BN;ReLU
64to256;resize back toinputgrid
Concat five256-channel results
1280 channels
Conv1×1;BatchNorm;ReLU;Dropout0.5
1280to64;dropout isidentity in eval
Two distinct1×1 deformable branches exist: aspp1 plus the first entry of parallel_block_sizes=(1,3,7).
Modulated deformable convolution
Offset Conv k×k
64to2k²:2,18,98 offsets for k1,3,7; biasTrue
Mask Conv k×k
64tok²:1,9,49 values;biasTrue
Multiply2 × Sigmoid
Modulation range(0,2)
Base kernel grid + learned offsets
Bilinear sample input features at displaced locations
Multiply sampled values by modulation and kernel weights
Learned kernel[256,64,k,k]; sum over channels/kernelpositions
Output256 channels
Stride1; spatial size preserved
regular_conv supplies its weights to torchvision deform_conv2d; it is not an extra executed convolution.
Image-reference primitives and values
Packed image tiles
Cin3072,768,192,48,3 fromcoarse tofull
Conv3×3
Cin to64;stride1,padding1
Conv3×3
64 toIPT outputchannels [384, 384, 192, 96, 48]
No normalization or activation between these two convolutions.
FeyNobg changes only Swin stage3 depth to24.
Training-only multiscale and gradient-label heads are stored
for checkpoint loading, but excluded from this eval graph.
Source: libreyolo/models/feynobg/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com