BEN2 Base

Click a block to read its description, or select it with Tab and Enter.

BEN2 BaseAlpha-matte logits,normalized RGB3 × 1024 × 1024,native eval. Five-field dimensions include thefield batch.LibreYOLOBEN2 BaseAlpha-matte logits,normalized RGB3 × 1024 × 1024,native eval. Five-field dimensions include thefield batch.Five fields and shared backboneNormalized RGB image3 × 1024 × 1024Four512² quadrants + global512² resize5 × 3 × 512 × 512;global bilinearresizeShared Swin-v1 backboneBasewidth128;depths[2,2,18,2];window12Five feature outputsPatch128×128²;stages128×128²,256×64²,512×32²,1024×16²Independent Conv3×3 + InstanceNorm + GELUChannels128,128,256,512,1024 each to128Each level keeps5fields. Projected names:E1,E2,E3,E4,E5.E1/E2 grids128²;E3 grid64²;E4 grid32²;E5 grid16².Independent shallowConv3×3 makes128×1024² original-image features.Multi-field decoderMultiFieldCrossAttention onE54local +1global;5×128×16²E4 + bilinear resize(E5);Refinement45×128×32²Conv3×3 +InstanceNorm +GELU128to128;32²E3 + bilinear resize(previous);Refinement35×128×64²Conv3×3 +InstanceNorm +GELU128to128;64²E2 + bilinear resize(previous);Refinement25×128×128²Conv3×3 +InstanceNorm +GELU128to128;128²E1 + bilinear resize(previous);Refinement15×128×128²Conv3×3 +InstanceNorm +GELU128to128;128²Every refinement updateslocal fields andadds them back intoglobal context.Field merge and full-resolution outputReassemble4localfields; addresizedglobal128 × 256 × 256ThreeConv3×3 mask-head operations128to384to384to128;IN+GELU afterfirsttwoAddbilinear-resized shallowfeature128 × 256 × 256Nearest×2;Conv3×3+IN+GELU128 × 512 × 512Addbilinear-resized shallowfeature128 × 512 × 512Nearest×2;Conv3×3+IN+GELU128 × 1024 × 1024Conv3×3,padding1128to1;1 × 1024 × 1024 logitsTraining sideout1...5 arestored butdo notexecute ininference.Swin-v1 backbone internalsConvpatch4,stride4 +LayerNorm3to128;128²;retainthispatchfeatureSwinblocks,n=2128channels,4heads;stage1output128²PatchMergingConcat2×2:512;LN;Linear512to256;64²Swinblocks,n=2256channels,8heads;stage2output64²PatchMergingConcat2×2:1024;LN;Linear1024to512;32²Swinblocks,n=18512channels,16heads;stage3output32²PatchMergingConcat2×2:2048;LN;Linear2048to1024;16²Swinblocks,n=21024channels,32heads;stage4output16²Eachstage output getsits ownLayerNorm beforedecoder projection.MultiFieldCrossAttention at16²Reassemblelocal4×16² into32²128channelsAdaptivepools to16²,4²,2²;concat tokens276memorytokens,128channelsGlobalquery cross-attention256queries;K276+2D sinepositions,V276;1head128Residual+LN;FFN128to256to128;residual+LNGlobalstate16²Splitupdatedglobal into4quadrantsEachglobalquadrant8²=64memorytokensFourindependent localcross-attentionsEachlocal256queries toitsglobal64keys;1head128Residual+LN;FFN128to256to128;residual+LNLocal4×128×16²Concatlocal andglobal alongfield batch5×128×16²Sinepositions only onglobal Q/K. FFNs useGELU;dropout inactive ineval.MultiFieldRefinement at32²,64²,128²Split4local and1global fieldEach128channelsGlobalConv1×1(128to1)+sigmoidNearestresizeto2H×2H,split4tiles,gate localfeaturesSplitglobal into4quadrants;adaptivepoolsTargetH/2,H/4,H/8;concatper-fieldmemoryFourindependent cross-attentionsQ=H²;KV=336/1344/5376 forH32/64/128Localresidual+LNGatedlocalfeaturesprovide residualLinear128to256;GELU;Linear256to128Residual+LayerNormReassembleupdatedlocals;resizeandaddtoglobalNearestresizeof2H×2H toH×HConcatlocal/global5×128×H×H;returntokenattention alongsideThepublic decoder ignores auxiliarytokenattention outputs.Shared primitivesSwin block: LayerNorm,windowattention,residual,LayerNorm,Linear4C,GELU,LinearC,residual.C=[128,256,512,1024];MLP=[512,1024,2048,4096];headwidth32;window12,alternatingshift0/6.Q/K/V linear projections128to128 infieldattention;SwinQKV widths384,768,1536,3072ScaledQK transposeFieldattention headwidth128;Swinheadwidth32Addpositions orrelativewindowbias whereusedFieldglobal:sinepositions;Swin:529-entryrelativebiastable +shiftmaskSoftmax overkeysWeights×V;headconcat;outputprojectionOutputhasquery token countandoriginalchannelwidthConv3×3,padding1DeclaredCin/Cout;biasTrueInstanceNorm2dPerimagechannel spatialnormalizationGELUFourier/foreground-colour refinement isnot partofBEN2 Base.Source: libreyolo/models/ben2/nn.py and model.py. Revision a4d0ecc9e17f.libreyolo.com