DDColor T

Click a block to read its description, or select it with Tab and Enter.

DDColor TImage colorization,input3 × 512 × 512 grayscale RGB,native eval. Output is2-channel Lab chroma.LibreYOLODDColor TImage colorization,input3 × 512 × 512 grayscale RGB,native eval. Output is2-channel Lab chroma.ConvNeXt encoderImageNet normalization3 × 512 × 512Conv4×4,s4 thenLayerNorm3 to96;grid128²ConvNeXt block,n=396 channelsOutput LayerNorm: E196 × 128²LayerNorm thenConv2×2,s296 to192;grid64²ConvNeXt block,n=3192 channelsOutput LayerNorm: E2192 × 64²LayerNorm thenConv2×2,s2192 to384;grid32²ConvNeXt block,n=9384 channelsOutput LayerNorm: E3384 × 32²LayerNorm thenConv2×2,s2384 to768;grid16²ConvNeXt block,n=3768 channelsOutput LayerNorm: E4768 × 16²A pooled classification feature is computed but ignored by DDColor.Wide UNet and pixel featuresUnetBlockWideInput768;skip384;output512 × 32²From E4 + E3UnetBlockWideInput512;skip192;output512 × 64²From previous + E2UnetBlockWideInput512;skip96;output256 × 128²From previous + E1PixelShuffleICNR ×4256to4096conv;output256×512²;blurColor transformer memories:512×32²,512×64²,256×128².Color queries and chroma outputLearned query content andposition100 × 256 each;3 learned level embeddingsColor decoder layer,n=9Memory levels cycle0,1,2 three timesFinal LayerNorm100 × 256Color embedding MLP256to256to256to256;ReLU afterfirsttwoDot queries with pixel features[100,256] × [256,512,512] =100 color mapsConcat color maps and normalized input100+3=103 channels at512²Spectral-normalized Conv1×1103to2;biasTrue;noactivationLab a,b prediction2 × 512 × 512;no sigmoid oroutputdenormalizationPostprocessing combines resized a,b with the original LabL.ConvNeXt blockInput featureStage widths [96, 192, 384, 768]DepthwiseConv7×7,padding3Groups equalstagewidth;biasTrueChannels-last LayerNormepsilon1e-6Linear expansionOutputwidths [384, 768, 1536, 3072]GELULinear projectionRestore stagewidthLayerScaleLearned channelvector,init1e-6+UnetBlockWide and shuffleSpectral-normalized Conv1×1 +BNCi to4Co;Co512,512,256ReLU thenPixelShuffle2Co channels atdoubledgridReplicationPad(left1,top1);AvgPool2,s1Blur preserves expandedsizeConcat with BatchNorm(skip),thenReLUConcat widths512+384,512+192,256+96Spectral-normalized Conv3×3Concat toCo;s1,p1;biasTrueReLU thenBatchNormOutput512,512,256channelsFinal×4shuffle uses256to4096Conv1×1,ReLU,PixelShuffle4,blur;it has spectral normalization and no extraBatchNorm.One color decoder layerInput query state100 × 256Cross-attention,8 headsQ=query+querypos;K=memory+sinepos;V=memory+LayerNorm256 channelsSelf-attention,8 headsQ,K=query+querypos;V=query+LayerNorm256 channelsLinear FFN256to2048ReLULinear FFN2048to256+LayerNorm256 channelsCross-attention executes before self-attention; dropout0,post-norm.Attention, position and color-MLP primitivesSeparate Q/K/V Linear256to2568 heads,32channels/headQK transpose /sqrt32Cross-memory lengths1024,4096,16384;selflength100Softmax overkeysWeights × V;concatheads;Linear256to256Memory Conv1×1 projections:512to256,512to256,256to256.Each projected memory adds its learned256-channel levelvector.2D normalized sine/cosine positions provide256 channels.Linear256to256ReLULinear256to256ReLULinear256to256Source: libreyolo/models/ddcolor/nn.py and model.py. Revision a4d0ecc9e17f.libreyolo.com