OWLv2 b16
Fit diagram
Read at 100%
Clear selection
Download SVG
Download PNG
Click a block to read its description, or select it with Tab and Enter.
OWLv2 b16
Open-vocabulary detection, 960 × 960 image, 3 class prompts of 16 tokens. Tensor sizes exclude batch.
LibreYOLO
OWLv2 b16
Open-vocabulary detection, 960 × 960 image, 3 class prompts of 16 tokens. Tensor sizes exclude batch.
Vision tower
Image
3 × 960 × 960
Conv2d 16×16 / 16
768 channels, bias=False
Flatten patches
3600 × 768
Learned CLS
1 × 768
Concat CLS and patches
3601 × 768
Position embedding
3601 × 768
+
LayerNorm before encoder
3601 × 768
Vision encoder block
3601 × 768, n=12
Post LayerNorm
3601 × 768
Select CLS and broadcast
3600 × 768
Select patch tokens
3600 × 768
Multiply patch × CLS
3600 × 768
LayerNorm
3600 × 768
Text tower
Tokenized class prompts
3 × 16 IDs, vocabulary 49,408
Token embedding
3 × 16 × 512
Embedded sequences
3 × 16 × 512
Position embedding
16 × 512
+
Text encoder block
16 × 512 per class, n=12
Final LayerNorm
3 × 16 × 512
Select EOS position
Argmax token ID per class
Linear text projection
512 to 512, no bias
L2 normalize
3 × 512
Causal attention plus padding mask.
Query validity is first token ID > 0.
Per-patch class scoring
Image patch features
3600 × 768
Text embeddings
3 × 512
Linear class embedding
768 to 512
L2 normalize with epsilon
3600 × 512, eps=1e-6
L2 normalize with epsilon
3 × 512, eps=1e-6
Einsum cosine similarity
3600 × 3
Linear logit shift
768 to 1 per patch
Linear logit scale
768 to 1 per patch
+
ELU
3600 × 1
Add 1
3600 × 1
Multiply shifted logits by scale
3600 × 3
Mask invalid class prompts
Invalid scores become dtype minimum
Returned logits are per patch and per requested class.
Vision encoder block
Input 3601 × 768
LayerNorm
3601 × 768
Multi-head self-attention
12 heads
+
LayerNorm
3601 × 768
Linear
768 to 3072
QuickGELU
3601 × 3072
Linear
3072 to 768
+
Output 3601 × 768
Vision self-attention
Query input
3601 × 768
Key/value input
3601 × 768
Linear Q
768 to 768
Reshape heads
12 × 3601 × 64
Linear K
768 to 768
Reshape heads
12 × 3601 × 64
Linear V
768 to 768
Reshape heads
12 × 3601 × 64
MatMul Q K-transpose
12 × 3601 × 3601
Scale
Divide by sqrt(64)
Softmax over keys
12 × 3601 × 3601
MatMul attention × V
12 × 3601 × 64
Merge heads
3601 × 768
Linear output
768 to 768
Box prediction
Image patch features
3600 × 768
Linear
768 to 768
GELU
3600 × 768
Linear
768 to 768
GELU
3600 × 768
Linear
768 to 4
Add grid box bias
Logit normalized corner and patch-size priors
Sigmoid
3600 × 4 normalized cxcywh
All prediction-head Linear layers include bias.
Objectness prediction
Image patch features
3600 × 768
Linear
768 to 768
GELU
3600 × 768
Linear
768 to 768
GELU
3600 × 768
Linear
768 to 1
Squeeze last axis
3600 objectness logits
All prediction-head Linear layers include bias.
Text encoder block
Input 16 × 512
LayerNorm
16 × 512
Multi-head self-attention
8 heads
+
LayerNorm
16 × 512
Linear
512 to 2048
QuickGELU
16 × 2048
Linear
2048 to 512
+
Output 16 × 512
Text causal self-attention
Query input
16 × 512
Key/value input
16 × 512
Linear Q
512 to 512
Reshape heads
8 × 16 × 64
Linear K
512 to 512
Reshape heads
8 × 16 × 64
Linear V
512 to 512
Reshape heads
8 × 16 × 64
MatMul Q K-transpose
8 × 16 × 16
Scale
Divide by sqrt(64)
Causal mask
16 × 16
+
Softmax over keys
8 × 16 × 16
MatMul attention × V
8 × 16 × 64
Merge heads
16 × 512
Linear output
512 to 512
QuickGELU
Input x
Elementwise activation
Multiply by 1.702
Same shape
Sigmoid
Same shape
Multiply by original x
Same shape
Detection outputs
Class logits
Sigmoid, max class per patch
Score threshold
Keep matching patch predictions
Convert normalized cxcywh to XYXY
Scale to the original image dimensions
Results
Boxes, confidence and requested class IDs
Objectness logits are a separate returned branch.
No DETR query decoder or FPN exists in this graph.
Variant values
Size
S / P
Np / Nv
Ev / nv / hv
Et / nt / ht
Mv / Mt
D
b16
960 / 16
3600 / 3601
768 / 12 / 12
512 / 12 / 8
3072 / 2048
512
l14
1008 / 14
5184 / 5185
1024 / 24 / 16
768 / 12 / 12
4096 / 3072
768
Nv includes CLS; Np is patch count. Ev/Et: tower widths; nv/nt: repeats; hv/ht: heads; Mv/Mt: MLP widths; D: text projection.
Source: libreyolo/models/owlv2/nn.py. Revision a4d0ecc9e17f.
libreyolo.com