按任务分类的模型

库中的每个系列,由 v1.5.0 模型注册表生成:17 类任务下共 82 个系列。

目标检测

40 个系列

为目标画出边界框,也是这个库最主要的任务。

RF-DETR

n, s, m, l for detection, pose and oriented boxes; n through xx for segmentation

YOLOv9

yolo9: t, s, m, c at 640 px; yolo9_p2: t, s at 640 px

D-FINE

n, s, m, l, x at 640 px

DEIM

deim: n, s, m, l, x at 640 px

EdgeCrafter

s, m, l, x at 640 px

RT-DETR

rtdetr: r18, r34, r50, r50m, r101, l, x at 640 px; rtdetrv2: r18, r34, r50, r50m, r101 for detection at 640 px, n, s, m, l, x for oriented boxes at 1024 px; rtdetrv4: s, m, l, x at 640 px

YOLO-NAS

s, m, l at 640 px

Dome-DETR

s, m, l at 800 px

PicoDet

s, m, l at 320 to 640 px

RTMDet

t, s, m, l, x at 640 px

YOLOv7

b at 640 px

YOLOX

n, t, s, m, l, x at 416 to 640 px

Florence-2

base, large at 768 px

Grounding DINO

t, b at 800 px

InternVL3

1b, 2b, 8b at 448 px

Kosmos-2

224 at 224 px

LFM2-VL

450m at 512 px

LibreMODUS
LocateAnything

3b at 2500 px

OMDet-Turbo

t at 640 px

OV-DEIM

s, m, l at 640 px

OWLv2

b16, l14 at 960 to 1008 px

Qwen3-VL

2b, 4b, 8b at 1024 px

SenseNova-Vision

7b at 1024 px

SmolVLM2

500m at 512 px

CenterNet

resdcn18, dla34 at 512 px

Deformable DETR

r50ss, r50ssdc5, r50, r50refine, r50twostage at 800 px

DETR

r50, r50dc5, r101, r101dc5 at 800 px

DINO-DETR

r50, r50s5, swinl at 800 px

EfficientDet
Faster R-CNN

n, s, m, l at 320 to 800 px

FCOS

r50 at 800 px

LW-DETR

t, s, m, l, x at 640 px

Mask R-CNN

r50 at 800 px

RetinaNet

r50, r50v2 at 800 px

SSD

300 at 300 px

YOLOv1

t, b at 448 px

YOLOv2

t, b at 416 to 608 px

YOLOv3

t, b, spp at 416 to 608 px

YOLOv4

t, b at 416 to 608 px

实例分割

13 个系列

为每个实例输出掩码,而不只是边界框。

RF-DETR

n, s, m, l for detection, pose and oriented boxes; n through xx for segmentation

D-FINE

n, s, m, l, x at 640 px

EdgeCrafter

s, m, l, x at 640 px

RTMDet

t, s, m, l, x at 640 px

MobileSAM

tiny at 1024 px

PicoSAM3

pico at 96 px

SAM

base, large, huge at 1024 px

SAM 3

large at 1008 px

SenseNova-Vision

7b at 1024 px

EoMT

s, b, l at 512 px

Mask R-CNN

r50 at 800 px

关键点

5 个系列

为每个检测到的实例输出骨架和关键点。

RF-DETR

n, s, m, l for detection, pose and oriented boxes; n through xx for segmentation

EdgeCrafter

s, m, l, x at 640 px

YOLO-NAS

s, m, l at 640 px

SenseNova-Vision

7b at 1024 px

图像分类

12 个系列

为整张图像给出一个标签,同时也是其他任务的主干网络。

ConvNeXt

t, s, b at 224 px

DINOv2

n, s, m, l at 518 px

EfficientNetV2

b0, b1, b2, b3 at 224 to 300 px

MobileNetV4

s, m, l at 224 to 256 px

ResNet

18, 34, 50, 101 at 224 px

AlexNet

b at 224 px

Swin Transformer

t, s, b, l at 224 px

VGG

16, 19, 16bn, 19bn at 224 px

ViT

ti, s, b, l at 224 px

DeiT

t, s, b at 224 px

旋转框

2 个系列

带角度的边界框,适用于航拍图像等非水平对齐的目标。

RF-DETR

n, s, m, l for detection, pose and oriented boxes; n through xx for segmentation

RT-DETR

rtdetr: r18, r34, r50, r50m, r101, l, x at 640 px; rtdetrv2: r18, r34, r50, r50m, r101 for detection at 640 px, n, s, m, l, x for oriented boxes at 1024 px; rtdetrv4: s, m, l, x at 640 px

深度估计

6 个系列

从单张图像估计每个像素的距离。

LibreMODUS
SenseNova-Vision

7b at 1024 px

Depth Anything 3

l at 504 px

Depth Anything V2

s, b, l, g at 518 px

MiDaS

s, l at 256 to 384 px

ZipDepth

b, bnpu at 384 px

语义分割

7 个系列

为每个像素分配类别,但不区分实例。

DINOv2

n, s, m, l at 518 px

LingBot-Vision
SegFormer
DeepLabv3
EoMT

s, b, l at 512 px

FCN

r50, r101 at 520 px

PIDNet

s, m, l at 1024 px

全景分割

2 个系列

一次前向同时输出语义和实例掩码。

SenseNova-Vision

7b at 1024 px

EoMT

s, b, l at 512 px

点检测与计数

3 个系列

用中心点代替边界框,轻到可以跑在单片机上。

LocateAnything

3b at 2500 px

SenseNova-Vision

7b at 1024 px

文字识别

2 个系列

在图像中定位并读出文字。

SenseNova-Vision

7b at 1024 px

PP-OCRv5

t, l at 960 px

特征向量

4 个系列

用于检索、聚类和相似度搜索的向量。

DINOv2

n, s, m, l at 518 px

LibreFaceRec

视线估计

1 个系列

判断人正在看向哪里。

L2CS-Net

r18, r34, r50, r101, r152 at 448 px

边缘检测

3 个系列

轮廓与边界。

LibreMODUS
DexiNed

b at 352 px

TEED

t at 352 px

表面法线

2 个系列

每个表面朝向哪个方向。

LibreMODUS
MoGe-2

s, b, l at 518 px

抠图

2 个系列

Alpha 通道抠图,连发丝和边缘一起保留。

BiRefNet

t, l at 1024 px

FeyNobg

l at 1024 px

图像修复

3 个系列

去噪、去模糊和超分辨率。

NAFNet

s, l at 256 px

Real-ESRGAN

x4, x2, x4t at 64 px

SwinIR

s, m, l at 64 px

三维网格

1 个系列

从图像重建三维几何。

SAM 3D Body

d3, h at 512 px