ZipDepth B
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
ZipDepth B
Relative inverse depth,input3 × 384 × 384,native unfused eval,global_mode=balanced. Shapes exclude batch.
LibreYOLO
ZipDepth B
Relative inverse depth,input3 × 384 × 384,native unfused eval,global_mode=balanced. Shapes exclude batch.
Encoder
RGB + ImageNet normalization
3 × 384 × 384
ConvBN 3×3,s2
3 to24;half skip24×192²
ConvBN3×3,s2
24 to48;48×96²
QARepBlock,n=2
48 to48;S1:48×96²
QARepBlock stride2
48 to96;48²
QARepBlock,n=2
96 to96;48²
MinimalMultiScale thenStripPooling
96×48²;S2
QARepBlock stride2
96 to192;24²
QARepBlock,n=6
192 to192;24²
ChannelAttention thenGlobalContext
192×24²;S3
QARepBlock stride2
192 to384;12²
QARepBlock,n=2
384×12²
LightweightSPPF
384×12²;S4
Cross-scale exchange and FPN
MinimalCrossScale(S3,S4)
S3:192×24², S4:384×12²
ConvBN1×1 on exchangedS4
384 to288;12²
UltraLightFusion3
high192/low288 to192;24²
UltraLightFusion2
high96/low192 to144;48²
UltraLightFusion1
high48/low144 to96;96²
UltraLightFusion half
high24/low96 to32;192²
High-resolution inputs come fromS3,S2,S1 andhalf-stem skip.
Half-resolution depth and final2× reconstruction
Half feature
32×192×192
Conv2d3×3
32 to1;s1,p1;halfdepth192²
Convex upsample
b:9-neighbor convex. bnpu:learnedinterpolation alpha.
ReLU
Nonnegative relative inverse depth
Output
1×384×384
Both trained heads also read the32-channel half feature.
b andbnpu have different head parameters; no weight-free rewrite.
Public loaded models fuse QARep branches; this view is unfused.
QARepBlock
Stage widths48,96,192,384. Transitions48to96,96to192,192to384.
Conv2d3×3
Ci toCo;declaredstride;biasFalse
BatchNorm2d
Co channels
Conv2d1×1
Ci toCo;declaredstride;biasFalse
BatchNorm2d
Co channels
+
+
Input identity
Only Ci=Co,s1
ReLU
One fusedConv3×3 can replacebranches afterload
Identity branch has no BatchNorm; downsample blocks omit it.
ChannelAttention
Input
192×24²
Mean overH,W
192×1×1
Conv1×1
192to24;biasFalse
ReLU
Conv1×1
24to192;biasFalse
Sigmoid
×
StripPoolingAttention
Mean overwidth
96×48×1
Mean overheight
96×1×48
+
DepthwiseConv1×1
96channels/groups;biasFalse
BatchNorm2d
96channels
Sigmoid
×
Input X
MinimalMultiScale
DepthwiseConv3×3
96groups;dilation1,padding1
DepthwiseConv3×3
96groups;dilation2,padding2
+
BatchNorm2d
96 channels
+
Input X
96×48²
Residual addition has no finalactivation.
GlobalContextBlock
Conv1×1;flatten
192to1;576spatial logits
Softmax over576positions
Context weights
Batch matmul X × weights
192×576 times576×1
Conv1×1
192to48
BatchNorm2d + ReLU
48channels
Conv1×1
48to192
+
Input X
LightweightSPPF
ConvBN1×1
384to96;12²
MaxPool5×5,stride1,padding2
96×12²
MaxPool5×5,stride1,padding2
96×12²
MaxPool5×5,stride1,padding2
96×12²
x
p1
p2
p3
Concat x,p1,p2,p3
384×12²
ConvBN1×1
384to384
Bidirectional MinimalCrossScale
GroupedConv1×1
384to192;groups4;biasFalse
Nearest resize12² to24²
Match targetgrid
Multiply0.3
+
OriginalS3
GroupedConv1×1
192to384;groups4;biasFalse
AvgPool2×2,stride2
Match targetgrid
Multiply0.3
+
OriginalS4
Both exchange paths read originalS3/S4, before either result is updated.
UltraLightFusion and ConvBN
High-resolution source
Channels192,96,48,24 across four fusions
Bilinear resize low-resolution source
Channels288,192,144,96;align_cornersFalse
GroupedConv1×1,groups4
Both project to192,144,96,32 respectively
GroupedConv1×1,groups4
Both project to192,144,96,32 respectively
+
BatchNorm2d thenReLU
192,144,96,32channels respectively
ConvBN elsewhere means Conv2d (biasFalse),BatchNorm2d,ReLU.
b:9-neighbor convex head
Conv3×3 +BN +ReLU
32to8;192²;biasFalse
Conv1×1
8to36 =9neighbors×4subpixels
Reshape andsoftmax over9neighbors
9×4×192×192;temperature1
Replicate-pad halfdepth;Unfold3×3
9×1×192×192 neighborhood values
Multiply masks × neighbors,sum9
4×192×192
PixelShuffle2
1×384×384
Depth
Both heads finish withReLU. They are separately trained checkpoint variants.
bnpu:unfold-free trained head
Conv1×1 +BN +ReLU
32to16;192²;biasFalse
DepthwiseConv5×5 +BN +ReLU
16groups;padding2;192²
Conv1×1;bilinear×2;sigmoid
16to1;alpha1×384×384
Nearest×2 andbilinear×2 halfdepth
Two1×384×384 candidate maps
alpha×nearest + (1-alpha)×bilinear
Learned convex interpolation
Depth
Both heads finish withReLU. They are separately trained checkpoint variants.
Balanced mode includes strip pooling andGC context; it does not execute the optional full-mode global-token attention.
Source: libreyolo/models/zipdepth/nn.py and model.py. Revision a4d0ecc9e17f.
libreyolo.com