BEN2 Base
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
BEN2 Base
Alpha-matte logits,normalized RGB3 × 1024 × 1024,native eval. Five-field dimensions include thefield batch.
LibreYOLO
BEN2 Base
Alpha-matte logits,normalized RGB3 × 1024 × 1024,native eval. Five-field dimensions include thefield batch.
Five fields and shared backbone
Normalized RGB image
3 × 1024 × 1024
Four512² quadrants + global512² resize
5 × 3 × 512 × 512;global bilinearresize
Shared Swin-v1 backbone
Basewidth128;depths[2,2,18,2];window12
Five feature outputs
Patch128×128²;stages128×128²,256×64²,512×32²,1024×16²
Independent Conv3×3 + InstanceNorm + GELU
Channels128,128,256,512,1024 each to128
Each level keeps5fields. Projected names:E1,E2,E3,E4,E5.
E1/E2 grids128²;E3 grid64²;E4 grid32²;E5 grid16².
Independent shallowConv3×3 makes128×1024² original-image features.
Multi-field decoder
MultiFieldCrossAttention onE5
4local +1global;5×128×16²
E4 + bilinear resize(E5);Refinement4
5×128×32²
Conv3×3 +InstanceNorm +GELU
128to128;32²
E3 + bilinear resize(previous);Refinement3
5×128×64²
Conv3×3 +InstanceNorm +GELU
128to128;64²
E2 + bilinear resize(previous);Refinement2
5×128×128²
Conv3×3 +InstanceNorm +GELU
128to128;128²
E1 + bilinear resize(previous);Refinement1
5×128×128²
Conv3×3 +InstanceNorm +GELU
128to128;128²
Every refinement updateslocal fields andadds them back intoglobal context.
Field merge and full-resolution output
Reassemble4localfields; addresizedglobal
128 × 256 × 256
ThreeConv3×3 mask-head operations
128to384to384to128;IN+GELU afterfirsttwo
Addbilinear-resized shallowfeature
128 × 256 × 256
Nearest×2;Conv3×3+IN+GELU
128 × 512 × 512
Addbilinear-resized shallowfeature
128 × 512 × 512
Nearest×2;Conv3×3+IN+GELU
128 × 1024 × 1024
Conv3×3,padding1
128to1;1 × 1024 × 1024 logits
Training sideout1...5 arestored butdo notexecute ininference.
Swin-v1 backbone internals
Convpatch4,stride4 +LayerNorm
3to128;128²;retainthispatchfeature
Swinblocks,n=2
128channels,4heads;stage1output128²
PatchMerging
Concat2×2:512;LN;Linear512to256;64²
Swinblocks,n=2
256channels,8heads;stage2output64²
PatchMerging
Concat2×2:1024;LN;Linear1024to512;32²
Swinblocks,n=18
512channels,16heads;stage3output32²
PatchMerging
Concat2×2:2048;LN;Linear2048to1024;16²
Swinblocks,n=2
1024channels,32heads;stage4output16²
Eachstage output getsits ownLayerNorm beforedecoder projection.
MultiFieldCrossAttention at16²
Reassemblelocal4×16² into32²
128channels
Adaptivepools to16²,4²,2²;concat tokens
276memorytokens,128channels
Globalquery cross-attention
256queries;K276+2D sinepositions,V276;1head128
Residual+LN;FFN128to256to128;residual+LN
Globalstate16²
Splitupdatedglobal into4quadrants
Eachglobalquadrant8²=64memorytokens
Fourindependent localcross-attentions
Eachlocal256queries toitsglobal64keys;1head128
Residual+LN;FFN128to256to128;residual+LN
Local4×128×16²
Concatlocal andglobal alongfield batch
5×128×16²
Sinepositions only onglobal Q/K. FFNs useGELU;dropout inactive ineval.
MultiFieldRefinement at32²,64²,128²
Split4local and1global field
Each128channels
GlobalConv1×1(128to1)+sigmoid
Nearestresizeto2H×2H,split4tiles,gate localfeatures
Splitglobal into4quadrants;adaptivepools
TargetH/2,H/4,H/8;concatper-fieldmemory
Fourindependent cross-attentions
Q=H²;KV=336/1344/5376 forH32/64/128
Localresidual+LN
Gatedlocalfeaturesprovide residual
Linear128to256;GELU;Linear256to128
Residual+LayerNorm
Reassembleupdatedlocals;resizeandaddtoglobal
Nearestresizeof2H×2H toH×H
Concatlocal/global
5×128×H×H;returntokenattention alongside
Thepublic decoder ignores auxiliarytokenattention outputs.
Shared primitives
Swin block: LayerNorm,windowattention,residual,LayerNorm,Linear4C,GELU,LinearC,residual.
C=[128,256,512,1024];MLP=[512,1024,2048,4096];headwidth32;window12,alternatingshift0/6.
Q/K/V linear projections
128to128 infieldattention;SwinQKV widths384,768,1536,3072
ScaledQK transpose
Fieldattention headwidth128;Swinheadwidth32
Addpositions orrelativewindowbias whereused
Fieldglobal:sinepositions;Swin:529-entryrelativebiastable +shiftmask
Softmax overkeys
Weights×V;headconcat;outputprojection
Outputhasquery token countandoriginalchannelwidth
Conv3×3,padding1
DeclaredCin/Cout;biasTrue
InstanceNorm2d
Perimagechannel spatialnormalization
GELU
Fourier/foreground-colour refinement isnot partofBEN2 Base.
Source: libreyolo/models/ben2/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com