Swin t

Click a block to read its description, or select it with Tab and Enter.

Swin tClassification, 224 × 224 input, 1,000 classes. Stage tensors use H × W × C, excluding batch.LibreYOLOSwin tClassification, 224 × 224 input, 1,000 classes. Stage tensors use H × W × C, excluding batch.NetworkInput NCHW3 × 224 × 224Conv2d 4×4 / 496 × 56 × 56Permute to NHWC56 × 56 × 96LayerNorm56 × 56 × 96Stage 156 × 56 × 96, n=2Stage 228 × 28 × 192, n=2Stage 314 × 14 × 384, n=6Stage 47 × 7 × 768, n=2LayerNorm7 × 7 × 768Mean over H and W768Linear classifier768 to 1,000 logitsWindow size: 7 × 7 throughout.Final stage: shift is disabled.No CLS token or absolute position.Dropout and stochastic depth are absent.Stage 1Input from patch embedding56 × 56 × 96Identity downsample56 × 56 × 96Window block, total repeats 2Odd block shift 0; even block shift 3.LayerNorm56 × 56 × 96Cyclic roll H and WOdd: 0; even: -3 pixelsPartition 7×7 windows64 windows, 49 × 96 eachWindow self-attention3 heads, width 32Reverse window partition56 × 56 × 96Reverse cyclic rollOdd: 0; even: +3 pixels+Flatten spatial tokens3136 × 96LayerNorm3136 × 96Linear96 to 384GELU3136 × 384Linear384 to 96+Reshape to NHWC56 × 56 × 96No spatial padding needed at this input.Stage 1 window attentionQuery input49 × 96Key/value input49 × 96Linear Q96 to 96Reshape heads3 × 49 × 32Linear K96 to 96Reshape heads3 × 49 × 32Linear V96 to 96Reshape heads3 × 49 × 32MatMul Q K-transpose3 × 49 × 49ScaleDivide by sqrt(32)Position bias3 × 49 × 49+Shift mask0 / -100, even blocks+Softmax over keys3 × 49 × 49MatMul attention × V3 × 49 × 32Merge heads49 × 96Linear output96 to 96Stage 2Input from previous stage56 × 56 × 96Reshape 2×2 neighborhoods28 × 28 × 384LayerNorm384 channelsLinear reduction (no bias)384 to 192Window block, total repeats 2Odd block shift 0; even block shift 3.LayerNorm28 × 28 × 192Cyclic roll H and WOdd: 0; even: -3 pixelsPartition 7×7 windows16 windows, 49 × 192 eachWindow self-attention6 heads, width 32Reverse window partition28 × 28 × 192Reverse cyclic rollOdd: 0; even: +3 pixels+Flatten spatial tokens784 × 192LayerNorm784 × 192Linear192 to 768GELU784 × 768Linear768 to 192+Reshape to NHWC28 × 28 × 192No spatial padding needed at this input.Stage 2 window attentionQuery input49 × 192Key/value input49 × 192Linear Q192 to 192Reshape heads6 × 49 × 32Linear K192 to 192Reshape heads6 × 49 × 32Linear V192 to 192Reshape heads6 × 49 × 32MatMul Q K-transpose6 × 49 × 49ScaleDivide by sqrt(32)Position bias6 × 49 × 49+Shift mask0 / -100, even blocks+Softmax over keys6 × 49 × 49MatMul attention × V6 × 49 × 32Merge heads49 × 192Linear output192 to 192Stage 3Input from previous stage28 × 28 × 192Reshape 2×2 neighborhoods14 × 14 × 768LayerNorm768 channelsLinear reduction (no bias)768 to 384Window block, total repeats 6Odd block shift 0; even block shift 3.LayerNorm14 × 14 × 384Cyclic roll H and WOdd: 0; even: -3 pixelsPartition 7×7 windows4 windows, 49 × 384 eachWindow self-attention12 heads, width 32Reverse window partition14 × 14 × 384Reverse cyclic rollOdd: 0; even: +3 pixels+Flatten spatial tokens196 × 384LayerNorm196 × 384Linear384 to 1536GELU196 × 1536Linear1536 to 384+Reshape to NHWC14 × 14 × 384No spatial padding needed at this input.Stage 3 window attentionQuery input49 × 384Key/value input49 × 384Linear Q384 to 384Reshape heads12 × 49 × 32Linear K384 to 384Reshape heads12 × 49 × 32Linear V384 to 384Reshape heads12 × 49 × 32MatMul Q K-transpose12 × 49 × 49ScaleDivide by sqrt(32)Position bias12 × 49 × 49+Shift mask0 / -100, even blocks+Softmax over keys12 × 49 × 49MatMul attention × V12 × 49 × 32Merge heads49 × 384Linear output384 to 384Stage 4Input from previous stage14 × 14 × 384Reshape 2×2 neighborhoods7 × 7 × 1536LayerNorm1536 channelsLinear reduction (no bias)1536 to 768Window block, total repeats 2All blocks use shift 0.LayerNorm7 × 7 × 768Cyclic roll H and W0 pixelsPartition 7×7 windows1 windows, 49 × 768 eachWindow self-attention24 heads, width 32Reverse window partition7 × 7 × 768Reverse cyclic roll0 pixels+Flatten spatial tokens49 × 768LayerNorm49 × 768Linear768 to 3072GELU49 × 3072Linear3072 to 768+Reshape to NHWC7 × 7 × 768No spatial padding needed at this input.Stage 4 window attentionQuery input49 × 768Key/value input49 × 768Linear Q768 to 768Reshape heads24 × 49 × 32Linear K768 to 768Reshape heads24 × 49 × 32Linear V768 to 768Reshape heads24 × 49 × 32MatMul Q K-transpose24 × 49 × 49ScaleDivide by sqrt(32)Position bias24 × 49 × 49+Shift mask0 only+Softmax over keys24 × 49 × 49MatMul attention × V24 × 49 × 32Merge heads49 × 768Linear output768 to 768Variant valuesSizeC1, C2, C3, C4n1, n2, n3, n4h1, h2, h3, h4t96, 192, 384, 7682, 2, 6, 23, 6, 12, 24s96, 192, 384, 7682, 2, 18, 23, 6, 12, 24b128, 256, 512, 10242, 2, 18, 24, 8, 16, 32l192, 384, 768, 15362, 2, 18, 26, 12, 24, 48Source: libreyolo/models/swin/classifier.py. Revision a4d0ecc9e17f.libreyolo.com