Swin t
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
Swin t
Classification, 224 × 224 input, 1,000 classes. Stage tensors use H × W × C, excluding batch.
LibreYOLO
Swin t
Classification, 224 × 224 input, 1,000 classes. Stage tensors use H × W × C, excluding batch.
Network
Input NCHW
3 × 224 × 224
Conv2d 4×4 / 4
96 × 56 × 56
Permute to NHWC
56 × 56 × 96
LayerNorm
56 × 56 × 96
Stage 1
56 × 56 × 96, n=2
Stage 2
28 × 28 × 192, n=2
Stage 3
14 × 14 × 384, n=6
Stage 4
7 × 7 × 768, n=2
LayerNorm
7 × 7 × 768
Mean over H and W
768
Linear classifier
768 to 1,000 logits
Window size: 7 × 7 throughout.
Final stage: shift is disabled.
No CLS token or absolute position.
Dropout and stochastic depth are absent.
Stage 1
Input from patch embedding
56 × 56 × 96
Identity downsample
56 × 56 × 96
Window block, total repeats 2
Odd block shift 0; even block shift 3.
LayerNorm
56 × 56 × 96
Cyclic roll H and W
Odd: 0; even: -3 pixels
Partition 7×7 windows
64 windows, 49 × 96 each
Window self-attention
3 heads, width 32
Reverse window partition
56 × 56 × 96
Reverse cyclic roll
Odd: 0; even: +3 pixels
+
Flatten spatial tokens
3136 × 96
LayerNorm
3136 × 96
Linear
96 to 384
GELU
3136 × 384
Linear
384 to 96
+
Reshape to NHWC
56 × 56 × 96
No spatial padding needed at this input.
Stage 1 window attention
Query input
49 × 96
Key/value input
49 × 96
Linear Q
96 to 96
Reshape heads
3 × 49 × 32
Linear K
96 to 96
Reshape heads
3 × 49 × 32
Linear V
96 to 96
Reshape heads
3 × 49 × 32
MatMul Q K-transpose
3 × 49 × 49
Scale
Divide by sqrt(32)
Position bias
3 × 49 × 49
+
Shift mask
0 / -100, even blocks
+
Softmax over keys
3 × 49 × 49
MatMul attention × V
3 × 49 × 32
Merge heads
49 × 96
Linear output
96 to 96
Stage 2
Input from previous stage
56 × 56 × 96
Reshape 2×2 neighborhoods
28 × 28 × 384
LayerNorm
384 channels
Linear reduction (no bias)
384 to 192
Window block, total repeats 2
Odd block shift 0; even block shift 3.
LayerNorm
28 × 28 × 192
Cyclic roll H and W
Odd: 0; even: -3 pixels
Partition 7×7 windows
16 windows, 49 × 192 each
Window self-attention
6 heads, width 32
Reverse window partition
28 × 28 × 192
Reverse cyclic roll
Odd: 0; even: +3 pixels
+
Flatten spatial tokens
784 × 192
LayerNorm
784 × 192
Linear
192 to 768
GELU
784 × 768
Linear
768 to 192
+
Reshape to NHWC
28 × 28 × 192
No spatial padding needed at this input.
Stage 2 window attention
Query input
49 × 192
Key/value input
49 × 192
Linear Q
192 to 192
Reshape heads
6 × 49 × 32
Linear K
192 to 192
Reshape heads
6 × 49 × 32
Linear V
192 to 192
Reshape heads
6 × 49 × 32
MatMul Q K-transpose
6 × 49 × 49
Scale
Divide by sqrt(32)
Position bias
6 × 49 × 49
+
Shift mask
0 / -100, even blocks
+
Softmax over keys
6 × 49 × 49
MatMul attention × V
6 × 49 × 32
Merge heads
49 × 192
Linear output
192 to 192
Stage 3
Input from previous stage
28 × 28 × 192
Reshape 2×2 neighborhoods
14 × 14 × 768
LayerNorm
768 channels
Linear reduction (no bias)
768 to 384
Window block, total repeats 6
Odd block shift 0; even block shift 3.
LayerNorm
14 × 14 × 384
Cyclic roll H and W
Odd: 0; even: -3 pixels
Partition 7×7 windows
4 windows, 49 × 384 each
Window self-attention
12 heads, width 32
Reverse window partition
14 × 14 × 384
Reverse cyclic roll
Odd: 0; even: +3 pixels
+
Flatten spatial tokens
196 × 384
LayerNorm
196 × 384
Linear
384 to 1536
GELU
196 × 1536
Linear
1536 to 384
+
Reshape to NHWC
14 × 14 × 384
No spatial padding needed at this input.
Stage 3 window attention
Query input
49 × 384
Key/value input
49 × 384
Linear Q
384 to 384
Reshape heads
12 × 49 × 32
Linear K
384 to 384
Reshape heads
12 × 49 × 32
Linear V
384 to 384
Reshape heads
12 × 49 × 32
MatMul Q K-transpose
12 × 49 × 49
Scale
Divide by sqrt(32)
Position bias
12 × 49 × 49
+
Shift mask
0 / -100, even blocks
+
Softmax over keys
12 × 49 × 49
MatMul attention × V
12 × 49 × 32
Merge heads
49 × 384
Linear output
384 to 384
Stage 4
Input from previous stage
14 × 14 × 384
Reshape 2×2 neighborhoods
7 × 7 × 1536
LayerNorm
1536 channels
Linear reduction (no bias)
1536 to 768
Window block, total repeats 2
All blocks use shift 0.
LayerNorm
7 × 7 × 768
Cyclic roll H and W
0 pixels
Partition 7×7 windows
1 windows, 49 × 768 each
Window self-attention
24 heads, width 32
Reverse window partition
7 × 7 × 768
Reverse cyclic roll
0 pixels
+
Flatten spatial tokens
49 × 768
LayerNorm
49 × 768
Linear
768 to 3072
GELU
49 × 3072
Linear
3072 to 768
+
Reshape to NHWC
7 × 7 × 768
No spatial padding needed at this input.
Stage 4 window attention
Query input
49 × 768
Key/value input
49 × 768
Linear Q
768 to 768
Reshape heads
24 × 49 × 32
Linear K
768 to 768
Reshape heads
24 × 49 × 32
Linear V
768 to 768
Reshape heads
24 × 49 × 32
MatMul Q K-transpose
24 × 49 × 49
Scale
Divide by sqrt(32)
Position bias
24 × 49 × 49
+
Shift mask
0 only
+
Softmax over keys
24 × 49 × 49
MatMul attention × V
24 × 49 × 32
Merge heads
49 × 768
Linear output
768 to 768
Variant values
Size
C1, C2, C3, C4
n1, n2, n3, n4
h1, h2, h3, h4
t
96, 192, 384, 768
2, 2, 6, 2
3, 6, 12, 24
s
96, 192, 384, 768
2, 2, 18, 2
3, 6, 12, 24
b
128, 256, 512, 1024
2, 2, 18, 2
4, 8, 16, 32
l
192, 384, 768, 1536
2, 2, 18, 2
6, 12, 24, 48
Source: libreyolo/models/swin/classifier.py. Revision a4d0ecc9e17f.
libreyolo.com