DDColor T
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
DDColor T
Image colorization,input3 × 512 × 512 grayscale RGB,native eval. Output is2-channel Lab chroma.
LibreYOLO
DDColor T
Image colorization,input3 × 512 × 512 grayscale RGB,native eval. Output is2-channel Lab chroma.
ConvNeXt encoder
ImageNet normalization
3 × 512 × 512
Conv4×4,s4 thenLayerNorm
3 to96;grid128²
ConvNeXt block,n=3
96 channels
Output LayerNorm: E1
96 × 128²
LayerNorm thenConv2×2,s2
96 to192;grid64²
ConvNeXt block,n=3
192 channels
Output LayerNorm: E2
192 × 64²
LayerNorm thenConv2×2,s2
192 to384;grid32²
ConvNeXt block,n=9
384 channels
Output LayerNorm: E3
384 × 32²
LayerNorm thenConv2×2,s2
384 to768;grid16²
ConvNeXt block,n=3
768 channels
Output LayerNorm: E4
768 × 16²
A pooled classification feature is computed but ignored by DDColor.
Wide UNet and pixel features
UnetBlockWide
Input768;skip384;output512 × 32²
From E4 + E3
UnetBlockWide
Input512;skip192;output512 × 64²
From previous + E2
UnetBlockWide
Input512;skip96;output256 × 128²
From previous + E1
PixelShuffleICNR ×4
256to4096conv;output256×512²;blur
Color transformer memories:512×32²,512×64²,256×128².
Color queries and chroma output
Learned query content andposition
100 × 256 each;3 learned level embeddings
Color decoder layer,n=9
Memory levels cycle0,1,2 three times
Final LayerNorm
100 × 256
Color embedding MLP
256to256to256to256;ReLU afterfirsttwo
Dot queries with pixel features
[100,256] × [256,512,512] =100 color maps
Concat color maps and normalized input
100+3=103 channels at512²
Spectral-normalized Conv1×1
103to2;biasTrue;noactivation
Lab a,b prediction
2 × 512 × 512;no sigmoid oroutputdenormalization
Postprocessing combines resized a,b with the original LabL.
ConvNeXt block
Input feature
Stage widths [96, 192, 384, 768]
DepthwiseConv7×7,padding3
Groups equalstagewidth;biasTrue
Channels-last LayerNorm
epsilon1e-6
Linear expansion
Outputwidths [384, 768, 1536, 3072]
GELU
Linear projection
Restore stagewidth
LayerScale
Learned channelvector,init1e-6
+
UnetBlockWide and shuffle
Spectral-normalized Conv1×1 +BN
Ci to4Co;Co512,512,256
ReLU thenPixelShuffle2
Co channels atdoubledgrid
ReplicationPad(left1,top1);AvgPool2,s1
Blur preserves expandedsize
Concat with BatchNorm(skip),thenReLU
Concat widths512+384,512+192,256+96
Spectral-normalized Conv3×3
Concat toCo;s1,p1;biasTrue
ReLU thenBatchNorm
Output512,512,256channels
Final×4shuffle uses256to4096Conv1×1,ReLU,PixelShuffle4,blur;
it has spectral normalization and no extraBatchNorm.
One color decoder layer
Input query state
100 × 256
Cross-attention,8 heads
Q=query+querypos;K=memory+sinepos;V=memory
+
LayerNorm
256 channels
Self-attention,8 heads
Q,K=query+querypos;V=query
+
LayerNorm
256 channels
Linear FFN
256to2048
ReLU
Linear FFN
2048to256
+
LayerNorm
256 channels
Cross-attention executes before self-attention; dropout0,post-norm.
Attention, position and color-MLP primitives
Separate Q/K/V Linear256to256
8 heads,32channels/head
QK transpose /sqrt32
Cross-memory lengths1024,4096,16384;selflength100
Softmax overkeys
Weights × V;concatheads;Linear256to256
Memory Conv1×1 projections:512to256,512to256,256to256.
Each projected memory adds its learned256-channel levelvector.
2D normalized sine/cosine positions provide256 channels.
Linear256to256
ReLU
Linear256to256
ReLU
Linear256to256
Source: libreyolo/models/ddcolor/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com