CLIP b32 classify

Click a block to read its description, or select it with Tab and Enter.

CLIP b32 classifyImage input 224 × 224; text length 77. Classify example uses 3 classes and one prompt per class.LibreYOLOCLIP b32 classifyImage input 224 × 224; text length 77. Classify example uses 3 classes and one prompt per class.Image towerImage3 × 224 × 224Conv2d 32×32 / 32768 × 7 × 7, no biasFlatten and transpose49 × 768Learned CLS1 × 768Concat CLS and patches50 × 768Position embedding50 × 768+LayerNorm before blocks50 × 768Vision transformer block50 × 768, n=12LayerNorm after blocks50 × 768Select CLS768MatMul learned projection768 to 512, no biasL2 normalize512Text towerText token IDs77 IDs, vocabulary 49,408Token embedding77 × 512Embedded sequence77 × 512Position embedding77 × 512+Text transformer block77 × 512, n=12LayerNorm77 × 512Select EOT positionArgmax token ID, width 512MatMul learned projection512 to 512, no biasL2 normalize512Task outputNormalized image512Normalized class matrix3 × 512MatMul cosine similarities3 class scoresMultiply exp(logit_scale)3 logitsPrediction postprocess applies softmax across the configured class set.Normalized image: task outputNormalized text: task outputVision transformer blockInput 50 × 768LayerNorm50 × 768Multi-head self-attention12 heads+LayerNorm50 × 768Linear768 to 3072GELU50 × 3072Linear3072 to 768+Output 50 × 768Vision self-attentionQuery input50 × 768Key/value input50 × 768Linear Q768 to 768Reshape heads12 × 50 × 64Linear K768 to 768Reshape heads12 × 50 × 64Linear V768 to 768Reshape heads12 × 50 × 64MatMul Q K-transpose12 × 50 × 50ScaleDivide by sqrt(64)Softmax over keys12 × 50 × 50MatMul attention × V12 × 50 × 64Merge heads50 × 768Linear output768 to 768Text transformer blockInput 77 × 512LayerNorm77 × 512Multi-head self-attention8 heads+LayerNorm77 × 512Linear512 to 2048GELU77 × 2048Linear2048 to 512+Output 77 × 512Text causal self-attentionQuery input77 × 512Key/value input77 × 512Linear Q512 to 512Reshape heads8 × 77 × 64Linear K512 to 512Reshape heads8 × 77 × 64Linear V512 to 512Reshape heads8 × 77 × 64MatMul Q K-transpose8 × 77 × 77ScaleDivide by sqrt(64)Causal mask77 × 77+Softmax over keys8 × 77 × 77MatMul attention × V8 × 77 × 64Merge heads77 × 512Linear output512 to 512Variant valuesSizePEvnvhvEtnthtb32327681212512128b16167681212512128l1414102424167681212D: 512 for b32/b16, 768 for l14. Nv: 50, 197, 257 respectively.Ev/Et: vision/text width. nv/nt: repeats. hv/ht: heads. P: patch size.Concrete views contain all resolved widths, head counts and token counts.Source: libreyolo/models/clip/nn.py. Revision a4d0ecc9e17f.libreyolo.com