These are the frozen docs for v1.3.0. They are kept for anyone pinned to that release and are no longer updated. Read the docs for the current release.
这是早期版本 v1.3.0 的文档。最新文档为 v1.4.0,可在版本菜单中找到。
两份配套指南更深入地介绍专门主题: LibreVLM 指南 介绍封装 Qwen3-VL、Florence-2 等视觉语言模型的开放词表检测; 实验性任务指南 介绍更多实验性工作流,包括 LoRA / DoRA 微调。分类、旋转框和姿态也在下文各自的章节中有文档说明。
简介
v1.3.0 验证范围
经过充分测试的路径是 YOLO9 和 RF-DETR 的检测、训练与推理,包括 RF-DETR 分割。
其他模型系列、任务以及多 GPU 工作流均可使用,但仍属实验性。
LibreYOLO 是一个采用 MIT 许可的计算机视觉工具包。v1.3.0 提供了涵盖检测、分割、分类、深度等的广泛模型目录,但其经过验证的支持范围是有意收窄的:
- YOLO9 检测 - CNN 路径。
- RF-DETR 检测 - transformer 路径。
- RF-DETR 分割 - 经过充分测试的分割路径。
我们建议将这些路径作为新项目的默认选择,因为它们在检测、训练与推理方面经过了最充分的测试。其他受支持的系列和任务通过同一个统一的 LibreYOLO() 工厂工作,但在 v1.3.0 中属于实验性。如果你有特定理由,可以使用它们。
1 from libreyolo import LibreYOLO, SAMPLE_IMAGE 2 3 # Default: YOLO9 detection 4 model = LibreYOLO("LibreYOLO9c.pt") 5 result = model(SAMPLE_IMAGE, conf=0.25, save=True) 6 7 print(f"Detected {len(result)} objects") 8 print(result.boxes.xyxy) 9 print(result.saved_path)
核心特性
- 对 YOLO9 检测、RF-DETR 检测和 RF-DETR 分割进行了充分测试并作为推荐默认项
- 统一的
LibreYOLO()工厂,用于加载检查点、导出产物和运行时 - 通过一致的 API 完成检测、分割、姿态和视线任务
- 支持图像、目录和视频推理(大尺寸帧可选分块推理)
- 通过 ByteTrack 内置多目标跟踪
- ONNX、TorchScript、TensorRT、OpenVINO、NCNN 和 CoreML 导出,内嵌元数据,并配有相应的运行时后端
- 兼容 COCO 的验证,提供 mAP 指标,并包含分割和姿态验证器
- Ultralytics 风格的
libreyolo命令行工具,用于 predict / train / val / export - 接受任意图像格式:文件路径、URL、PIL、NumPy、PyTorch 张量、原始字节
v1.3.0 中的重大变更
- DAMO-YOLO 已移除,且无别名:
LibreDAMOYOLO会抛出AttributeError,加载 DAMO-YOLO 检查点时会被拒绝。 - YOLO9 仅支持检测。
-seg、-pose、-cls、-obb和-sem的 YOLO9 变体已移除。 - RF-DETR 移除了 classify、semantic 和 depth。 它现在的任务是 detect、segment、pose 和 obb。分类和语义迁移到新的
LibreDINOv2系列;深度迁移到新的LibreDepthAnythingV2系列。 - TFLite 导出需要 Python 3.12+(onnx2tf wheel 约束)。onnx2tf 的最低版本提高到
>=2.4.3,并移除了旧的 onnx2tf 运行时补丁。
兼容性
可将此矩阵作为 v1.3.0 的快速支持速查表。✓ 表示已验证路径,exp 表示实验性, prev 表示研究预览,空白单元格表示当前不支持。只有 YOLO9 和 RF-DETR 检测(外加 RF-DETR 分割)经过充分测试;其余全部,包括新增的分类、语义、深度和点定位系列,均为实验性。
| 模型系列 | v1.3.0 状态 | 推理 | 训练 | 检测 | 分割 | 语义 | 分类 | 姿态 | OBB | 深度 | 点 | 视线 | ONNX | TorchScript | TensorRT | OpenVINO | NCNN | CoreML | TFLite |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| YOLO9 | 已验证 detect,单 GPU | ✓ | ✓ | ✓ | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | exp |
| RF-DETR | 已验证 detect + segment;pose / OBB 预览 | ✓ | ✓ | ✓ | ✓ | Not currently supported | Not currently supported | prev | prev | Not currently supported | Not currently supported | Not currently supported | ✓ | ✓ | ✓ | ✓ | Not currently supported | exp | exp |
| YOLOX | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | exp | exp | Not currently supported |
| YOLO9-E2E | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported |
| YOLO-NAS | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | exp | Not currently supported | Not currently supported |
| D-FINE | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | exp | Not currently supported | Not currently supported |
| DEIM | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | exp | Not currently supported | Not currently supported |
| DEIMv2 | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | exp | Not currently supported | Not currently supported |
| RT-DETR | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | exp | exp | exp | Not currently supported |
| RT-DETRv2 | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| RT-DETRv4 | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| PicoDet | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| RTMDet | 实验性 | exp | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| EC | 实验性 | exp | exp | exp | exp | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| DINOv2 | 新增,实验性(需要 transformers) | exp | exp | Not currently supported | Not currently supported | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| MobileNetV4 | 新增,实验性分类器(Apache) | exp | exp | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| ConvNeXt | 新增,实验性分类器(Apache) | exp | exp | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| EfficientNetV2 | 新增,实验性分类器(Apache) | exp | exp | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| Depth Anything V2 | 新增,实验性;不支持导出 | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| FOMO | 新增,实验性;不自动下载 | exp | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
| L2CS | 实验性,仅推理 | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | exp | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported | Not currently supported |
Depth Anything V2 没有导出路径。TFLite 导出为实验性,仅限 YOLO9 检测和 RF-DETR detect / segment / pose。分类系列(MobileNetV4、ConvNeXt、EfficientNetV2)支持 ONNX 导出。CoreML 导出会生成 .mlpackage 包并需要 libreyolo[coreml]:仅限 macOS、不支持 INT8, 且对 RF-DETR、D-FINE、DEIM、DEIMv2 或 EC 不内嵌 NMS。
安装
环境要求
- Python 3.10+
- PyTorch 1.13+ 和 torchvision 0.11+
从 PyPI 安装
1 pip install libreyolo
本文档对应即将发布的 v1.3.0。在 v1.3.0 发布到 PyPI 之前,请使用源码安装来获取本页所述的功能。
从源码安装
1 git clone https://github.com/LibreYOLO/libreyolo.git 2 cd libreyolo 3 git checkout dev 4 pip install -e .
可选依赖
1 # ONNX export and inference 2 pip install libreyolo[onnx] 3 # or: pip install onnx onnxsim onnxruntime 4 5 # RT-DETR compatibility extra (currently no extra packages) 6 pip install libreyolo[rtdetr] 7 8 # RF-DETR support 9 pip install libreyolo[rfdetr] 10 # or: pip install transformers 11 12 # TensorRT export and inference (NVIDIA GPU) 13 pip install libreyolo[tensorrt] 14 # Installs TensorRT CUDA 12 Python packages on Linux/Windows. 15 # Host driver/CUDA compatibility still matters. 16 17 # OpenVINO export and inference (Intel CPU/GPU/VPU) 18 pip install libreyolo[openvino] 19 # INT8 export also needs: pip install nncf 20 21 # NCNN export and inference 22 pip install libreyolo[ncnn] 23 # or: pip install pnnx ncnn 24 25 # ByteTrack API compatibility extra 26 pip install libreyolo[tracking] 27 # Tracking dependencies are part of the base install in v1.3.0. 28 29 # CoreML export and inference (macOS only for runtime) 30 pip install libreyolo[coreml] 31 # or: pip install coremltools 32 33 # L2CS gaze optional auto-download helper 34 pip install libreyolo[gaze] 35 # Optional parity with the upstream RetinaFace-based L2CS pipeline 36 pip install libreyolo[gaze-retinaface] 37 38 # Install every optional LibreYOLO extra 39 pip install libreyolo[all]
如果使用 uv,最可靠的做法是为每个 extra 建立独立的 venv:
1 # ONNX environment 2 uv venv .venv-onnx 3 uv pip install --python .venv-onnx/bin/python -e '.[onnx]' 4 5 # RT-DETR environment 6 uv venv .venv-rtdetr 7 uv pip install --python .venv-rtdetr/bin/python -e '.[rtdetr]' 8 9 # Repeat with .[rfdetr], .[openvino], .[ncnn], .[coreml], .[gaze], .[tracking], or .[tensorrt] as needed
这样可以避免改动项目环境,并保持可选依赖相互隔离。TensorRT、OpenVINO、NCNN 和 CoreML 等特定厂商 extra 可能仍需平台相关的原生包。
快速开始
若要使用测试最充分的路径,请选择单 GPU 的 YOLO9 检测、RF-DETR 检测或 RF-DETR 分割。它们通过同一个工厂加载、接受相同的输入并返回相同的 Results 对象,因此你可以在它们之间切换而无需改动周边代码。
YOLO9 - CNN 旗舰
1 from libreyolo import LibreYOLO, SAMPLE_IMAGE 2 3 # Use the official checkpoint name and let the factory resolve the details 4 model = LibreYOLO("LibreYOLO9c.pt") 5 6 # Run on a single image (SAMPLE_IMAGE ships with the package) 7 result = model(SAMPLE_IMAGE) 8 9 print(f"Found {len(result)} objects") 10 print(result.boxes.xyxy) # bounding boxes (N, 4) 11 print(result.boxes.conf) # confidence scores (N,) 12 print(result.boxes.cls) # class IDs (N,)
RF-DETR - transformer 旗舰
1 from libreyolo import LibreYOLO, SAMPLE_IMAGE 2 3 # Same factory, same call shape - just point at an RF-DETR checkpoint 4 model = LibreYOLO("LibreRFDETRs.pt") 5 result = model(SAMPLE_IMAGE) 6 7 print(f"Found {len(result)} objects") 8 print(result.boxes.xyxy)
保存标注后的输出
1 result = model(SAMPLE_IMAGE, save=True) 2 print(result.saved_path) # e.g. runs/detect/predict/parkour.jpg
处理目录
1 results = model("images/", save=True, batch=4) 2 for r in results: 3 print(f"{r.path}: {len(r)} detections")
可用模型
推荐的已验证路径:YOLO9 检测或 RF-DETR 检测 / 分割
这些模型的检测、训练与推理经过最充分的测试。在 v1.3.0 中,请将其他系列、任务以及多 GPU 工作流视为实验性。
LibreYOLO v1.3.0 提供两个已验证的旗舰系列,以及更广泛的受支持和新增模型目录。每个模型都通过同一个 LibreYOLO() 工厂加载,但只有下面的已验证路径才应被视为经过充分测试。
YOLO9 - CNN 旗舰
Recommended| 尺寸 | 代号 | 输入尺寸 | 适用场景 | 检测检查点 |
|---|---|---|---|---|
| Tiny | "t" | 640 | 快速推理 | LibreYOLO9t.pt |
| Small | "s" | 640 | 均衡 | LibreYOLO9s.pt |
| Medium | "m" | 640 | 更高精度 | LibreYOLO9m.pt |
| Compact | "c" | 640 | 最佳精度 | LibreYOLO9c.pt |
在 v1.3.0 中 YOLO9 仅支持检测。非检测的旗舰变体(包括旧的 -seg 检查点)已移除;如需分割,请使用下文的 RF-DETR 或 EdgeCrafter。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreYOLO9c.pt") # detection
RF-DETR - transformer 旗舰
Recommended| 尺寸 | 代号 | 输入尺寸 | 适用场景 | 检测检查点 |
|---|---|---|---|---|
| Nano | "n" | 384 | 边缘端 | LibreRFDETRn.pt |
| Small | "s" | 512 | 均衡 | LibreRFDETRs.pt |
| Medium | "m" | 576 | 更高精度 | LibreRFDETRm.pt |
| Large | "l" | 704 | 最高精度 | LibreRFDETRl.pt |
LibreYOLO 在 Hugging Face 组织上提供 Apache 干净授权的 RF-DETR 检测尺寸 N/S/M/L。XL/2XL 档位有意不提供。
充分测试 分割: LibreRFDETRn-seg.pt, LibreRFDETRs-seg.pt, LibreRFDETRm-seg.pt, LibreRFDETRl-seg.pt。 更大的 -seg 尺寸(x、xx) 会拉取上游 RF-DETR seg-XL / seg-2XL 权重,采用非商业许可, 且不托管在 LibreYOLO 组织上。请参见 分割 章节。
研究预览 姿态: LibreRFDETRx-pose.pt(移植自 RF-DETR v1.8.0 的 GroupPose;仅提供 576 下的 x 尺寸)。 OBB: LibreRFDETRn-obb.pt, LibreRFDETRs-obb.pt, LibreRFDETRm-obb.pt, LibreRFDETRl-obb.pt (旋转框,使用检测的输入尺寸)。请将两者视为研究预览,而非已验证路径。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreRFDETRs.pt") # detect (validated) 4 # model = LibreYOLO("LibreRFDETRs-seg.pt") # segment (validated) 5 # model = LibreYOLO("LibreRFDETRx-pose.pt") # pose (research preview) 6 # model = LibreYOLO("LibreRFDETRn-obb.pt") # obb (research preview)
其他检测系列
与已验证路径共享同一工厂和 API 接口的可检测系列。它们在 v1.3.0 中为实验性。每个检查点名称都链接到 LibreYOLO 组织上的模型卡; 将任意名称传给 LibreYOLO(),工厂会在首次使用时自动获取。
托管说明: YOLO-NAS 检查点(上方纯文本)托管在 Deci 的 CDN 上,采用其专有权重许可, 而非 LibreYOLO 的 Hugging Face 组织。工厂仍会在首次使用时自动下载它们。 DAMO-YOLO 已在 v1.3.0 中移除,不再可加载。
v1.3.0 中的新模型系列
v1.3.0 新增了分类、密集语义分割、单目深度和点定位系列。它们通过同一个工厂加载,但属于新增且实验性。DINOv2 需要 pip install libreyolo[rfdetr](transformers)。
| 系列 | 状态 | 任务 | 检查点 |
|---|---|---|---|
| MobileNetV4 | 实验性 | classify | LibreMobileNetV4s-cls.pt, LibreMobileNetV4m-cls.pt, LibreMobileNetV4l-cls.pt |
| ConvNeXt | 实验性 | classify | LibreConvNeXtt-cls.pt, LibreConvNeXts-cls.pt, LibreConvNeXtb-cls.pt |
| EfficientNetV2 | 实验性 | classify | LibreEfficientNetV2b0-cls.pt, LibreEfficientNetV2b1-cls.pt, LibreEfficientNetV2b2-cls.pt, LibreEfficientNetV2b3-cls.pt |
| DINOv2 | 实验性 | semantic, classify | LibreDINOv2n.pt, LibreDINOv2s.pt, LibreDINOv2m.pt, LibreDINOv2l.pt, LibreDINOv2n-cls.pt, LibreDINOv2s-cls.pt, LibreDINOv2m-cls.pt, LibreDINOv2l-cls.pt |
| Depth Anything V2 | 实验性 | depth | LibreDepthAnythingV2s-depth.pt, LibreDepthAnythingV2b-depth.pt, LibreDepthAnythingV2l-depth.pt, LibreDepthAnythingV2g-depth.pt |
| FOMO | 实验性 | point | LibreFOMOs-point.pt, LibreFOMOm-point.pt, LibreFOMOl-point.pt |
- MobileNetV4 是商用干净的分类路径:Apache-2.0 ImageNet-1k 权重(s/m/l 分别为 224/224/256),支持预测、top-1/top-5 验证、微调训练和 ONNX 导出。
- ConvNeXt(V1 Tiny/Small/Base,224)和 EfficientNetV2(b0-b3,224-300)是另外的 Apache-2.0 ImageNet-1k 分类器(精度档)。
- DINOv2 是带任务头的 DINOv2 主干:默认在 518 下进行密集语义分割,并在 224 下提供分类线性探针。它不是 RF-DETR 检测器。分类在 v1.3.0 中从 RF-DETR 迁移到这里。
- Depth Anything V2 进行单目深度估计(尺寸 s/b/l/g,均在 518)。ViT-S 权重为 Apache-2.0;ViT-B/L/G 为 CC-BY-NC-4.0(非商业)。仅支持推理和零样本验证:不可训练,且不支持导出。
- FOMO 是点定位器,为每个目标输出
(x, y, class, confidence)。预训练权重不再分发:请传入本地检查点或从头训练。
可提示与 VLM 档位: LibreSAM(可提示分割,libreyolo[sam]) 和 LibreVLM 视觉语言检测器档位 (libreyolo[vlm])是单独的类别,它们加载 上游的 Hugging Face 快照,且不通过 LibreYOLO() 检测器工厂路由。它们的权重继承 各上游模型的许可。
专用模型
| 系列 | 状态 | 任务 | 检查点 |
|---|---|---|---|
| L2CS | 实验性 | gaze(仅推理)- 参见 视线估计 | LibreL2CSr50.pt |
L2CS 架构尺寸包括 r18、r34、r50、r101 和 r152,但上游发布的 Gaze360 检查点为 ResNet-50。安装 libreyolo[gaze] 以获取可选的下载助手, 或为其他尺寸传入本地检查点路径。L2CS 权重不由 LibreYOLO 托管(Gaze360 数据集 许可禁止再分发)。
工厂函数
对每个模型和运行时都使用 LibreYOLO() 工厂。给它一个官方检查点名称或导出产物路径, 然后让它选择正确的模型系列、任务、类别数和运行时:
1 from libreyolo import LibreYOLO 2 3 # Default: YOLO9 detection 4 model = LibreYOLO("LibreYOLO9c.pt") 5 6 # Flagship transformer: RF-DETR 7 model = LibreYOLO("LibreRFDETRs.pt") 8 model = LibreYOLO("LibreRFDETRs-seg.pt") # validated segmentation 9 10 # New in v1.3.0 11 model = LibreYOLO("LibreMobileNetV4s-cls.pt") # classification (Apache, ImageNet-1k) 12 model = LibreYOLO("LibreDINOv2n.pt") # semantic segmentation 13 model = LibreYOLO("LibreDepthAnythingV2s-depth.pt") # monocular depth 14 model = LibreYOLO("LibreFOMOs-point.pt") # point localization (local weights) 15 16 # Exported deployment formats 17 model = LibreYOLO("model.onnx") # ONNX Runtime 18 model = LibreYOLO("model.engine") # TensorRT 19 model = LibreYOLO("model.mlpackage") # CoreML (macOS) 20 model = LibreYOLO("model_openvino/") # OpenVINO (directory) 21 model = LibreYOLO("model_ncnn/") # NCNN (directory)
对于可识别的官方检查点文件名,LibreYOLO 可以自动下载缺失的权重。对于自定义文件名,请指向明确的本地路径。 新项目请保持使用 YOLO9 检测或 RF-DETR 检测 / 分割; 其他系列、任务以及新增系列在 v1.3.0 中为实验性。
任务与文件名
LibreYOLO 使用统一的文件名约定,因此工厂仅凭检查点名称即可识别系列、尺寸和任务:
1 Libre<FAMILY><size>[-<task>].pt
任务后缀
| 任务 | 规范名称 | 文件名后缀 | 所属系列 |
|---|---|---|---|
| 检测 | "detect" | (无 - 隐式) | 大多数系列(默认) |
| 实例分割 | "segment" | -seg | RF-DETR、EdgeCrafter |
| 语义分割 | "semantic" | -sem | DINOv2 |
| 姿态估计 | "pose" | -pose | YOLO-NAS、EdgeCrafter、RF-DETR(预览) |
| 旋转框 | "obb" | -obb | RF-DETR(预览) |
| 分类 | "classify" | -cls | MobileNetV4、ConvNeXt、EfficientNetV2、DINOv2 |
| 单目深度 | "depth" | -depth | Depth Anything V2 |
| 点定位 | "point" | -point | FOMO |
| 视线估计 | "gaze" | -gaze | L2CS |
检测是隐式的(无后缀),遵循常见的 YOLO 约定。 工厂在 API 边界接受别名 ("detection"、"seg"、"keypoints"、"cls" 等); 只有上面的规范名称才会出现在文件名中。只有当某个任务在该系列的 受支持任务集中时,它才可用。
解析优先级
加载模型时,任务按以下顺序解析:
1 explicit task= → checkpoint["task"] → filename suffix → family default
1 from libreyolo import LibreYOLO 2 3 # 1. Filename suffix decides → segment 4 model = LibreYOLO("LibreRFDETRs-seg.pt") 5 6 # 2. Override regardless of filename 7 model = LibreYOLO("custom_weights.pt", task="segment") 8 9 # 3. Detection is implicit 10 model = LibreYOLO("LibreYOLO9c.pt") # task="detect"
各系列任务支持
| 系列 | v1.3.0 状态 | 默认 | 受支持任务 |
|---|---|---|---|
| YOLO9 | 单 GPU detect 充分测试;多 GPU 实验性 | detect | detect |
| RF-DETR | 单 GPU detect 和 segment 充分测试;pose 和 OBB 研究预览 | detect | detect, segment, pose, obb |
| YOLOX | 实验性 | detect | detect |
| YOLO9-E2E | 实验性 | detect | detect |
| YOLO-NAS | 实验性 | detect | detect, pose |
| D-FINE / DEIM / DEIMv2 | 实验性 | detect | detect |
| RT-DETR / RT-DETRv2 / RT-DETRv4 | 实验性 | detect | detect |
| PicoDet / RTMDet | 实验性 | detect | detect |
| EdgeCrafter (EC) | 实验性 | detect | detect, pose, segment |
| DINOv2 | 新增,实验性 | semantic | semantic, classify |
| MobileNetV4 / ConvNeXt / EfficientNetV2 | 新增,实验性 | classify | classify |
| Depth Anything V2 | 新增,实验性 | depth | depth |
| FOMO | 新增,实验性 | point | point |
| L2CS | 实验性 | gaze | gaze(仅推理) |
示例
1 # Detection (implicit) 2 LibreYOLO9c.pt 3 LibreRFDETRs.pt 4 LibreRTDETRr50.pt 5 6 # Instance segmentation (-seg) 7 LibreRFDETRs-seg.pt 8 LibreECm-seg.pt 9 10 # Semantic segmentation (-sem) 11 LibreDINOv2n.pt # semantic is DINOv2's default; -sem optional 12 13 # Pose (-pose) 14 LibreYOLONASn-pose.pt 15 LibreECs-pose.pt 16 LibreRFDETRx-pose.pt # preview; size x only 17 18 # Oriented boxes (-obb) 19 LibreRFDETRn-obb.pt # preview 20 21 # Classification (-cls) 22 LibreMobileNetV4s-cls.pt 23 LibreConvNeXtt-cls.pt 24 LibreEfficientNetV2b0-cls.pt 25 LibreDINOv2n-cls.pt # DINOv2 linear probe 26 27 # Depth (-depth) 28 LibreDepthAnythingV2s-depth.pt 29 30 # Point (-point) 31 LibreFOMOs-point.pt 32 33 # Gaze (-gaze optional; only task for L2CS) 34 LibreL2CSr50.pt
已弃用别名
LibreYOLORTDETR 和 LibreYOLORFDETR 分别是 LibreRTDETR 和 LibreRFDETR 的旧名称。它们仍可解析,但会发出 DeprecationWarning - 方便时请更新导入。
预测
单 GPU 预测路径在 YOLO9 检测、RF-DETR 检测和 RF-DETR 分割上经过充分测试。其他系列和任务使用相同的 API,但在 v1.3.0 中为实验性。
基础预测
1 result = model("image.jpg")
全部预测参数
1 result = model( 2 "image.jpg", 3 conf=0.25, # confidence threshold (default: 0.25) 4 iou=0.45, # NMS IoU threshold (default: 0.45) 5 imgsz=640, # input size override (default: model's native) 6 device="auto", # "auto", "cpu", "mps", "0", "cuda:0", ... 7 classes=[0, 2, 5], # filter to specific class IDs (default: all) 8 max_det=300, # max detections per image (default: 300) 9 augment=False, # test-time augmentation where implemented 10 save=True, # save annotated image (default: False) 11 batch=4, # directory batch size 12 stream=False, # video only: yield frame results instead of a list 13 vid_stride=1, # video only: process every N-th frame 14 show=False, # video only: display annotated frames 15 tiling=False, # large-image tiled detection 16 overlap_ratio=0.2, # tile overlap ratio 17 output_path="out/", # where to save (default: runs/detect/predict*/) 18 color_format="auto", # "auto", "rgb", or "bgr" 19 output_file_format="png", # output format: "jpg", "png", "webp" 20 )
model.predict(...) 是 model(...) 的别名。
支持的输入格式
LibreYOLO 接受以下任意格式的图像:
1 # File path (string or pathlib.Path) 2 result = model("photo.jpg") 3 result = model(Path("photo.jpg")) 4 5 # URL 6 result = model("https://example.com/image.jpg") 7 result = model("s3://bucket/image.jpg") 8 result = model("gs://bucket/image.jpg") 9 10 # PIL Image 11 from PIL import Image 12 img = Image.open("photo.jpg") 13 result = model(img) 14 15 # NumPy array (HWC or CHW, RGB or BGR, uint8 or float32) 16 import numpy as np 17 arr = np.random.randint(0, 255, (480, 640, 3), dtype=np.uint8) 18 result = model(arr) 19 20 # OpenCV (BGR) - specify color_format 21 import cv2 22 frame = cv2.imread("photo.jpg") 23 result = model(frame, color_format="bgr") 24 25 # PyTorch tensor (CHW or NCHW) 26 import torch 27 tensor = torch.randn(3, 640, 640) 28 result = model(tensor) 29 30 # Raw bytes 31 with open("photo.jpg", "rb") as f: 32 result = model(f.read()) 33 34 # BytesIO 35 from io import BytesIO 36 result = model(BytesIO(open("photo.jpg", "rb").read())) 37 38 # Directory of images 39 results = model("images/", batch=4)
处理结果
每次预测都返回一个 Results 对象(对目录则返回它们的列表):
1 result = model("image.jpg") 2 3 # Number of detections 4 len(result) # e.g., 5 5 6 # Bounding boxes in xyxy format (x1, y1, x2, y2) 7 result.boxes.xyxy # tensor of shape (N, 4) 8 9 # Bounding boxes in xywh format (center_x, center_y, width, height) 10 result.boxes.xywh # tensor of shape (N, 4) 11 12 # Confidence scores 13 result.boxes.conf # tensor of shape (N,) 14 15 # Class IDs 16 result.boxes.cls # tensor of shape (N,) 17 18 # Combined data: [x1, y1, x2, y2, conf, cls] 19 # Tracking adds a track_id column before conf/cls. 20 result.boxes.data # shape (N, 6), or (N, 7) when tracked 21 22 # Metadata 23 result.orig_shape # (height, width) of original image 24 result.path # source file path (or None) 25 result.names # {0: "person", 1: "bicycle", ...} 26 27 # Move to CPU / convert to numpy 28 result_cpu = result.cpu() 29 boxes_np = result.boxes.numpy()
类别过滤
将检测结果过滤到特定类别 ID:
1 # Only detect people (class 0) and cars (class 2) 2 result = model("image.jpg", classes=[0, 2])
内存中的批量推理
v1.3.0 新增:model.predict() 接受由内存中图像组成的列表或元组 (NumPy 数组、PIL 图像或张量),并将它们作为真正的堆叠前向批次运行。设置 batch > 1 可在支持的系列上 真正批处理前向计算;返回结果列表,每个输入对应一个。
1 import numpy as np 2 from libreyolo import LibreYOLO 3 4 model = LibreYOLO("LibreYOLO9c.pt") 5 6 frames = [ 7 np.random.randint(0, 255, (480, 640, 3), dtype=np.uint8), 8 np.random.randint(0, 255, (480, 640, 3), dtype=np.uint8), 9 np.random.randint(0, 255, (480, 640, 3), dtype=np.uint8), 10 ] 11 12 results = model(frames, batch=4) # list/tuple -> true batched inference 13 for r in results: 14 print(len(r), r.boxes.xyxy.shape)
模型信息
model.info()(v1.3.0 新增)返回一个 JSON 友好的字典,包含 系列、尺寸、任务、参数量、输入尺寸和类别名称,并在 verbose=True 时 打印可读的摘要。
1 meta = model.info(detailed=False, verbose=True) 2 # meta -> {"family": ..., "size": ..., "task": ..., "params": ..., "imgsz": ..., "names": {...}, ...}
分块推理
对于远大于模型输入尺寸的图像(如卫星影像、无人机画面),分块推理会将图像切分为重叠的分块,对每个分块运行检测,再合并结果。
在 v1.3.0 中分块仅支持检测。它会拒绝分割掩码,且不能与 augment=True 组合使用。
1 result = model( 2 "large_aerial_image.jpg", 3 tiling=True, 4 overlap_ratio=0.2, # 20% overlap between tiles (default) 5 save=True, 6 ) 7 8 # Extra metadata on tiled results 9 result.tiled # True 10 result.num_tiles # number of tiles used 11 result.saved_path # output directory when save=True 12 result.tiles_path # directory containing per-tile crops 13 result.grid_path # grid visualization image
在分块时设置 save=True,LibreYOLO 会保存:
final_image.jpg- 绘制了所有合并检测的完整图像grid_visualization.jpg- 显示分块网格叠加的图像tiles/- 各个分块裁剪图metadata.json- 分块参数和检测数量
如果图像已经小于模型的输入尺寸,则会自动跳过分块。
视频推理
将任意视频文件传给旗舰模型,LibreYOLO 会根据扩展名自动检测格式。支持:.mp4、.avi、.mov、.mkv、.webm、.gif 以及其他常见容器。
保存标注后的视频
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreYOLO9c.pt") 4 results = model("clip.mp4", save=True) 5 # Saved under runs/detect/predict*/clip.mp4
流式结果(内存平稳)
对于长视频,传入 stream=True 可获得一个生成器。每次迭代产出一帧的 Results - 不会在内存中缓存完整列表。
1 for result in model("long_clip.mp4", stream=True): 2 print(f"frame {result.frame_idx}: {len(result)} detections")
帧抽样
1 # Process every 2nd frame (halves compute and saved fps) 2 results = model("clip.mp4", vid_stride=2, save=True)
实时预览
1 # Display annotated frames in an OpenCV window while processing 2 results = model("clip.mp4", show=True)
用于自定义流水线的 VideoSource / VideoWriter
当你需要完全控制解码和编码时 - 自定义帧变换、混入跟踪输出、写入非默认编解码器 - 可直接使用这些构建块:
1 from libreyolo import LibreYOLO 2 from libreyolo.utils.video import VideoSource, VideoWriter 3 4 model = LibreYOLO("LibreYOLO9c.pt") 5 6 with VideoSource("clip.mp4", vid_stride=1) as src, \ 7 VideoWriter("out.mp4", fps=src.fps, width=src.width, height=src.height) as out: 8 for frame_bgr, frame_idx in src: 9 result = model(frame_bgr, color_format="bgr") 10 # ... draw, transform, etc. 11 out.write_frame(frame_bgr)
跟踪
LibreYOLO 提供两种运动跟踪器,它们消费任意检测器的 Results 并添加持久的轨迹 ID:ByteTrack(默认) 和 OC-SORT(v1.3.0 新增),后者对遮挡和 非线性运动更鲁棒。跟踪在单 GPU 的 YOLO9 检测和 RF-DETR 检测上测试最充分;其他检测系列在 v1.3.0 中为实验性。
安装
1 pip install libreyolo[tracking] # compatibility extra; tracking deps ship in base dev install
视频跟踪助手
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreYOLO9c.pt") 4 5 for result in model.track( 6 "clip.mp4", 7 track_conf=0.25, 8 iou=0.45, 9 save=True, # writes runs/track/<video_stem>.mp4 by default 10 vid_stride=1, 11 ): 12 print(result.frame_idx, result.track_id)
model.track() 是用于视频文件的生成器。它逐帧运行检测,内部使用较低的 ByteTrack 置信度进行恢复,并产出已填充 result.track_id 和 result.boxes.id 的 Results。
基础循环
1 from libreyolo import LibreYOLO, ByteTracker 2 from libreyolo.utils.video import VideoSource 3 4 model = LibreYOLO("LibreYOLO9c.pt") 5 tracker = ByteTracker() 6 7 with VideoSource("clip.mp4") as src: 8 for frame_bgr, frame_idx in src: 9 result = model(frame_bgr, color_format="bgr", conf=0.1) 10 tracked = tracker.update(result) 11 12 for i in range(len(tracked.boxes)): 13 track_id = int(tracked.boxes.id[i]) 14 xyxy = tracked.boxes.xyxy[i].tolist() 15 cls = int(tracked.boxes.cls[i]) 16 print(f"frame {frame_idx} - id {track_id} cls {cls} {xyxy}")
调用 tracker.update() 之后,result.boxes.id 保存轨迹 ID,且 result.boxes.is_track 为 True。
TrackConfig 参数
1 from libreyolo import ByteTracker, TrackConfig 2 3 cfg = TrackConfig( 4 track_high_thresh=0.25, # first-stage match threshold 5 track_low_thresh=0.1, # second-stage (low-conf recovery) 6 new_track_thresh=0.25, # minimum conf to start a new track 7 match_thresh=0.8, # IoU cost cutoff (stage 1) 8 match_thresh_low=0.5, # IoU cost cutoff (stage 2) 9 match_thresh_unconfirmed=0.7, # IoU cost cutoff for unconfirmed tracks 10 track_buffer=30, # frames to keep lost tracks before removal 11 frame_rate=30, # scales track_buffer 12 fuse_score=True, # multiply IoU by detection score 13 minimum_consecutive_frames=1, # frames to confirm a new track 14 ) 15 tracker = ByteTracker(config=cfg)
片段之间重置
1 tracker.reset() # clears tracked / lost / removed lists and the ID counter
OC-SORT(抗遮挡)
在 model.track() 上用 tracker="ocsort" 选择 OC-SORT。ByteTrack 仍为默认。使用 OC-SORT 时, track_conf 映射到跟踪器的 det_thresh(在 ByteTrack 中它映射到 track_high_thresh)。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreYOLO9c.pt") 4 5 for result in model.track( 6 "clip.mp4", 7 tracker="ocsort", # "bytetrack" (default) or "ocsort" 8 track_conf=0.25, # maps to OC-SORT det_thresh 9 iou=0.45, 10 save=True, 11 ): 12 print(result.frame_idx, result.track_id)
传入 OCSortConfig 以完全控制。提供配置实例会按类型选择跟踪器, 因此此时会忽略 tracker= 字符串。
1 from libreyolo import LibreYOLO, OCSortConfig 2 3 cfg = OCSortConfig( 4 det_thresh=0.25, # boxes above this drive association and spawn new tracks 5 max_age=30, # frames a track survives without an observation 6 min_hits=3, # consecutive hits before a track is reported 7 iou_threshold=0.3, # minimum IoU for a valid association 8 delta_t=3, # frame span used to estimate velocity direction 9 inertia=0.2, # weight of the velocity-direction (momentum) term 10 use_byte=False, # enable the BYTE low-score recovery pass 11 ) 12 13 model = LibreYOLO("LibreYOLO9c.pt") 14 for result in model.track("clip.mp4", tracker_config=cfg, save=True): 15 print(result.frame_idx, result.track_id)
分割
v1.3.0 验证范围
经过充分测试的路径是 YOLO9 和 RF-DETR 的检测、训练与推理,包括 RF-DETR 分割。
其他模型系列、任务以及多 GPU 工作流均可使用,但仍属实验性。
RF-DETR 分割是 v1.3.0 中的分割路径,也是经过充分测试的选项。EdgeCrafter(-seg)也提供分割头,但为实验性。YOLO9 不再提供分割头:自 v1.3.0 起它仅支持检测。
运行分割
1 from libreyolo import LibreYOLO 2 3 # RF-DETR segmentation, the heavily tested segmentation path 4 model = LibreYOLO("LibreRFDETRs-seg.pt") 5 result = model("photo.jpg") 6 7 # EdgeCrafter segmentation is also available but experimental 8 # model = LibreYOLO("LibreECs-seg.pt") 9 10 # Segmentation returns boxes + masks 11 print(result.boxes.xyxy) # bounding boxes (N, 4) 12 print(result.boxes.cls) # class IDs (N,) 13 print(result.masks.data.shape) # (N, H, W) tensor of binary masks
掩码表示
1 # Raw bitmasks 2 result.masks.data # tensor (N, H, W) - original image resolution 3 4 # Polygon contours (one ndarray of (M, 2) per instance) 5 result.masks.xy # absolute pixel coords 6 result.masks.xyn # normalized to [0, 1] 7 8 # Move / convert like Boxes 9 result.masks.cpu() 10 result.masks.numpy()
保存标注后的输出
save=True 会自动绘制框和半透明掩码叠加。
1 model("photo.jpg", save=True)
训练分割
RF-DETR 分割使用 RF-DETR 的 COCO 格式训练流水线,属于经过充分测试的单 GPU 范围。EdgeCrafter 分割训练可用,但为实验性。YOLO9 分割训练已在 v1.3.0 中移除。
姿态估计
姿态(人体关键点)估计可在 YOLO-NAS (-pose)、 EdgeCrafter (-pose) 以及 v1.3.0 新增的 RF-DETR (-pose) 预览上运行。每个姿态模型都是单类别 ("person"),具有 17 个 COCO 关键点。
运行姿态
1 from libreyolo import LibreYOLO 2 3 # YOLO-NAS pose 4 model = LibreYOLO("LibreYOLONASs-pose.pt") 5 result = model("people.jpg") 6 7 # EdgeCrafter pose 8 # model = LibreYOLO("LibreECs-pose.pt") 9 10 # Per-person bbox + 17 keypoints 11 print(result.boxes.xyxy) # person boxes (N, 4) 12 print(result.keypoints.xy.shape) # (N, 17, 2) pixel coordinates
预览 RF-DETR 姿态仅提供一个 x 尺寸的检查点: LibreRFDETRx-pose.pt。它在 v1.3.0 中是研究预览。
1 # RF-DETR pose preview (size x only) 2 model = LibreYOLO("LibreRFDETRx-pose.pt") 3 result = model("people.jpg") 4 print(result.keypoints.xy.shape) # (N, 17, 2)
关键点 API
1 result.keypoints.xy # (N, K, 2) absolute pixel coords 2 result.keypoints.xyn # (N, K, 2) normalized to [0, 1] 3 result.keypoints.conf # (N, K) per-keypoint confidence (None if model doesn't emit it) 4 result.keypoints.has_visible # (N, K) bool - conf > 0 5 6 result.keypoints.cpu() 7 result.keypoints.numpy()
保存标注后的输出
1 model("people.jpg", save=True) # draws boxes + skeleton
YOLO-NAS 支持姿态训练;EdgeCrafter 姿态目前仅推理。RF-DETR 姿态为预览(仅 x 尺寸)。YOLO9 仅检测,不提供姿态检查点。
视线估计
视线方向估计由 LibreL2CS 系列提供,它是 L2CS-Net 的移植,具有 ResNet 主干和两个角度分箱分类头。这是一个两阶段模型:上游的人脸检测器定位人脸,然后视线头以弧度预测每张人脸的 pitch 和 yaw。它在 v1.3.0 中仅推理且为实验性。
安装
1 pip install libreyolo[gaze] # optional Google Drive helper for Gaze360 weights
已发布的 L2CS ResNet-50 权重在 Gaze360 上训练,LibreYOLO 不做镜像。在没有可选助手的情况下,请传入本地检查点路径,或按照 LibreL2CS 打印的手动下载说明操作。
两阶段推理
1 from libreyolo import LibreYOLO 2 from libreyolo.models.l2cs.face import resolve_face_detector 3 4 # Gaze head 5 gaze = LibreYOLO("LibreL2CSr50.pt") 6 7 # Wire any LibreYOLO detector trained on faces 8 face = LibreYOLO("path/to/face-detector.pt") 9 gaze.face_detector = resolve_face_detector(face) 10 11 result = gaze("portrait.jpg") 12 print(result.boxes.xyxy) # face boxes 13 print(result.gaze.data) # (N, 2) tensor - pitch, yaw in radians
解码角度
1 import math 2 3 for i in range(len(result.gaze)): 4 pitch_rad, yaw_rad = result.gaze.data[i].tolist() 5 pitch_deg = pitch_rad * 180.0 / math.pi 6 yaw_deg = yaw_rad * 180.0 / math.pi 7 print(f"face {i}: pitch={pitch_deg:.1f} deg, yaw={yaw_deg:.1f} deg")
在命令行中:libreyolo predict model=LibreL2CSr50.pt source=portrait.jpg --face-detector path/to/face.pt。
分类
v1.3.0 新增:整图分类。提供两个系列,面向不同需求。LibreMobileNetV4 是生产级分类器(Apache-2.0 ImageNet-1k 权重,可导出为 ONNX)。LibreDINOv2 配合 task=classify 是 DINOv2 主干加线性探针,非常适合迁移学习,但其发布的权重为演示级,且暂不支持导出。分类是 v1.3.0 的新增任务,因此细节可能仍会变化。
| 系列 | 检查点 | 输入 | 权重 | 微调 | ONNX 导出 |
|---|---|---|---|---|---|
| LibreMobileNetV4 | LibreMobileNetV4{s,m,l}-cls.pt | 224 / 224 / 256 | Apache-2.0 ImageNet-1k(生产级) | 交叉熵 | 支持 |
| LibreDINOv2(classify) | LibreDINOv2{n,s,m,l}-cls.pt | 224 | Imagenette 演示级(10 类) | 线性探针 | 不支持 |
LibreMobileNetV4(生产级分类器)
原生 MobileNetV4-conv 移植(源自 timm),其 1000 类 ImageNet-1k 权重可逐比特一致地加载。尺寸 s / m 在 224 运行,l 在 256 运行。检查点:
加载并预测。单张图像返回一个 Results;直接从中读取 .probs(传入列表则返回列表)。
1 from libreyolo import LibreYOLO 2 3 # MobileNetV4-conv-Small, Apache-2.0 ImageNet-1k weights (auto-downloaded if missing) 4 model = LibreYOLO("LibreMobileNetV4s-cls.pt") 5 result = model("cat.jpg") # single image -> one Results 6 7 probs = result.probs # whole-image class vector, length = num classes 8 print(probs.top1, probs.top1conf) # top-1 class id (int) and its confidence 9 print(probs.top5, probs.top5conf) # 5 class ids and 5 confidences 10 print(result.names[probs.top1]) # human-readable class name
微调到自定义类别集(ImageFolder 布局)。分类头会自动按数据集类别数重建;ImageNet 预训练主干可干净地迁移。
1 from libreyolo import LibreMobileNetV4 2 3 model = LibreMobileNetV4(size="s") # ImageNet-pretrained backbone 4 model.train( 5 data="imagenette160", # known name, dataset root, or .zip URL 6 epochs=5, 7 batch=64, 8 lr0=1e-3, # AdamW + cosine, 1-epoch warmup 9 imgsz=224, 10 )
验证(top-1 / top-5 精度):
1 model = LibreYOLO("LibreMobileNetV4s-cls.pt") 2 metrics = model.val(data="imagenette160") 3 print(metrics["metrics/accuracy_top1"]) 4 print(metrics["metrics/accuracy_top5"])
导出为 ONNX(已验证与 eager 模式逐比特一致)。ONNX 图输出单个 logits 张量。
1 model = LibreYOLO("LibreMobileNetV4s-cls.pt") 2 path = model.export(format="onnx", imgsz=224) # single output: logits [batch, num_classes] 3 4 # Interop note: the ONNX output is RAW LOGITS, not softmaxed. The PyTorch 5 # predict path applies softmax for you; non-Python consumers must apply it 6 # themselves before reading probabilities.
LibreDINOv2 分类(线性探针 / 迁移)
冻结风格的 DINOv2-S 编码器加可训练线性头,在 224 运行。n / s / m / l 尺寸只控制投影头宽度:四者共享同一个 DINOv2-S 编码器,因此发布的检查点精度几乎相同。随附的 -cls 权重为演示级(在 Imagenette 上训练,10 类),因此请将该系列视为迁移学习选项,而非可直接替换的 1000 类分类器。检查点:
加载并预测(与 MobileNetV4 相同的 Probs 接口):
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreDINOv2s-cls.pt") # DINOv2-S backbone + linear probe (224) 4 result = model("springer.jpg") 5 print(result.probs.top1, result.probs.top1conf)
微调用于迁移。用 task="classify" 构建全新模型以获得全新分类头,或加载随附的 -cls 检查点继续训练。为获得最佳精度,请从随附检查点而非全新分类头开始微调,并保持默认的 lr=1e-4(更高的学习率收敛更差)。
1 from libreyolo import LibreDINOv2 2 3 # Fresh DINOv2 backbone + random linear head, sized to the dataset 4 model = LibreDINOv2(size="s", task="classify", nb_classes=3) 5 model.train(data="path/to/imagefolder", epochs=5, lr=1e-4, batch=4) 6 7 # Validate the same way (top-1 / top-5) 8 metrics = model.val(data="path/to/imagefolder") 9 print(metrics["metrics/accuracy_top1"])
LibreDINOv2 未实现导出。如果你需要可导出的分类器,请使用 LibreMobileNetV4。
1 model = LibreYOLO("LibreDINOv2s-cls.pt") 2 model.export(format="onnx") 3 # raises NotImplementedError: Export is not yet implemented for LibreDINOv2.
数据集布局(两个系列)
分类使用 ImageNet 风格的 ImageFolder 目录树(按文件夹组织,而非标签文件)。类别索引按文件夹名称排序分配。data= 接受数据集根目录、已知名称(如 imagenette160)或 .zip URL。
1 dataset_root/ 2 train/ # required; one subfolder per class 3 class_a/img001.jpg 4 class_a/img002.jpg 5 class_b/img003.jpg 6 val/ # required for validation; same class folders as train 7 class_a/img010.jpg 8 class_b/img011.jpg
Results.probs 参考
1 probs = result.probs # Probs payload, 1-D vector of length = num classes 2 probs.data # raw tensor / ndarray of class probabilities 3 probs.top1 # int - argmax class id 4 probs.top5 # list - 5 class ids, highest first 5 probs.top1conf # float - confidence of the top-1 class 6 probs.top5conf # 5 confidences, aligned with probs.top5
- MobileNetV4 权重为生产级(Apache-2.0 ImageNet-1k,逐比特一致加载)。DINOv2 分类权重为演示级(Imagenette,10 类)。
- v1.3.0 中没有 LibreRFDETR 分类器。分类已迁移到 LibreMobileNetV4 和 LibreDINOv2 系列;旧的 LibreRFDETR*-cls 检查点在加载时会被拒绝。
- 使用默认配方从头微调 DINOv2 在 Imagenette 上的 top-1 上限约为 0.93,低于随附的 0.976。从随附的 -cls 检查点开始微调以恢复精度。
- ONNX 分类输出为原始 logits。请在非 Python 消费方中自行应用 softmax。
- 预测单张图像返回一个 Results。直接读取 result.probs,或传入列表并对列表取索引:model(["a.jpg"])[0].probs。
深度估计
v1.3.0 新增:通过 LibreDepthAnythingV2 进行单目深度估计,它是 Depth Anything V2 的移植(DINOv2 编码器加 DPT 头,NeurIPS 2024)。它预测一张密集的相对逆深度图:值越大越靠近相机,不隐含任何度量单位。尺寸 s / b / l / g 对应 ViT-S / B / L / G,均在 518 运行。深度估计是 v1.3.0 的新增任务,仅支持推理和零样本验证:不支持训练,也不支持导出。
检查点。只有 ViT-S 检查点为 Apache-2.0 并自动托管:LibreDepthAnythingV2s-depth.pt。更大的编码器 LibreDepthAnythingV2b-depth.pt, LibreDepthAnythingV2l-depth.pt, LibreDepthAnythingV2g-depth.pt 为 CC-BY-NC-4.0,LibreYOLO 不做再分发;请使用 weights/convert_depth_anything_v2_weights.py 转换官方上游检查点。
运行深度估计
输入 imgsz 必须能被 14 整除(DINOv2 的 patch 网格)。深度图在原始图像画布上返回。
1 from libreyolo import LibreYOLO 2 3 # ViT-S encoder, Apache-2.0 weights (commercial use OK) 4 model = LibreYOLO("LibreDepthAnythingV2s-depth.pt") 5 result = model("street.jpg") 6 7 depth = result.depth_map # DepthMap payload, (H, W) float on the original canvas 8 print(depth.data.shape) # (H, W) 9 print(depth.min, depth.max, depth.mean) # relative inverse depth: higher = closer 10 norm = depth.normalized() # rescaled to [0, 1] over finite values
DepthMap API
1 depth = result.depth_map 2 depth.data # (H, W) float tensor / ndarray, relative inverse depth 3 depth.min # min over finite values 4 depth.max # max over finite values 5 depth.mean # mean over finite values 6 depth.normalized() # (H, W) rescaled to [0, 1]; non-finite pixels become 0 7 8 depth.cpu() 9 depth.numpy()
零样本验证
验证通过共享的深度验证器以零样本方式运行,并报告标准深度指标(AbsRel、RMSE 和 delta 阈值)。验证器会信箱填充到固定正方形并排除填充像素;由于预测使用 Depth Anything 原生的保持纵横比缩放,非正方形的验证指标是预测的有据近似。
1 metrics = model.val(data="depth_dataset.yaml") 2 print(metrics["metrics/abs_rel"]) # absolute relative error (lower is better) 3 print(metrics["metrics/rmse"]) # root mean squared error 4 print(metrics["metrics/delta1"]) # fraction within a 1.25x ratio (higher is better)
不支持
1 model.train(data="...") # raises NotImplementedError - DA V2 is inference + val only 2 model.export(format="onnx") # raises NotImplementedError - depth export is out of scope
- 许可是分开的:ViT-S(尺寸 s)权重为 Apache-2.0,可用于商业用途。ViT-B / ViT-L / ViT-G(尺寸 b / l / g)为 CC-BY-NC-4.0(非商业),LibreYOLO 不做再分发。
- 若用于商业用途,请坚持使用尺寸 s。
- 深度为相对逆深度,没有度量单位。如果你需要以米为单位,请自行标定。
- imgsz 必须能被 14 整除。批量预测被禁用,因为保持纵横比缩放会导致每张图像尺寸不一。
点定位
LibreFOMO 是 FOMO 风格的点定位器(尺寸 s / m / l),用于质心式检测:每个检测不是框,而是一个图像坐标。预测以 result.points 形式返回。预训练的 LibreFOMO 权重不会自动下载,因此请传入本地检查点路径(或从头训练,这属于实验性,需要 allow_experimental=True)。
1 from libreyolo import LibreYOLO 2 3 # LibreFOMO weights are not hosted by LibreYOLO - pass a local checkpoint 4 model = LibreYOLO("path/to/LibreFOMOm-point.pt") 5 result = model("scene.jpg") 6 7 points = result.points # Points payload, (N, 4) rows: x, y, class, confidence 8 print(points.xy) # (N, 2) absolute pixel coords 9 print(points.xyn) # (N, 2) normalized to [0, 1] 10 print(points.cls, points.conf)
训练
v1.3.0 验证范围
经过充分测试的路径是 YOLO9 和 RF-DETR 的检测、训练与推理,包括 RF-DETR 分割。
其他模型系列、任务以及多 GPU 工作流均可使用,但仍属实验性。
经过充分测试的训练路径是单 GPU 的 YOLO9 检测、RF-DETR 检测和 RF-DETR 分割。其他模型系列的训练器和多 GPU 工作流可用,但为实验性。在 v1.3.0 中 YOLO9 仅检测,因此没有 YOLO9 分割或姿态训练。
YOLO9 - CNN 旗舰训练
1 from libreyolo import LibreYOLO 2 3 # Fine-tune from a pretrained checkpoint (recommended) 4 model = LibreYOLO("LibreYOLO9c.pt") 5 6 results = model.train( 7 data="coco128.yaml", # path to data.yaml (required) 8 9 # Schedule 10 epochs=300, # default: 300 11 batch=16, 12 imgsz=640, 13 14 # Optimizer 15 lr0=0.01, # initial learning rate 16 optimizer="SGD", # "SGD", "Adam", "AdamW" 17 18 # System 19 device="0", # "" | "cpu" | "cuda" | "0" | "0,1" 20 workers=8, 21 seed=0, 22 23 # Output 24 project="runs/train", 25 name="yolo9_exp", 26 exist_ok=False, 27 28 # Training features 29 amp=True, # automatic mixed precision 30 patience=50, # early stopping patience 31 resume=False, # resume from loaded checkpoint 32 pretrained=True, # transfer-learning init (True, a path, or None) 33 cache="disk", # cache decoded images: False | True/"ram" | "disk" 34 freeze=10, # freeze first N groups, or a list of indices / module names 35 save_plots=True, # write final validation plots to the run dir 36 ) 37 38 print(f"Best mAP50-95: {results['best_mAP50_95']:.3f}") 39 print(f"Best checkpoint: {results['best_checkpoint']}")
训练完成后,模型实例会自动以最优权重重新加载,因此你可以立即调用 model(...)。freeze、cache、pretrained 和 save_plots 为 v1.3.0 新增,并在所有基于训练器的系列中受支持。
RF-DETR - transformer 旗舰训练
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreRFDETRs.pt") 4 5 results = model.train( 6 data="path/to/data.yaml", 7 epochs=100, 8 batch_size=4, # NOTE: RF-DETR uses batch_size, not batch 9 lr=1e-4, 10 output_dir="runs/train/rfdetr_exp", 11 )
RF-DETR 有自己的训练签名(batch_size、lr、output_dir),但共用 LibreYOLO 的数据集加载器。为检测或分割传入 data.yaml,可采用 YOLO TXT 或原生 COCO JSON 布局 —— 参见 数据集格式。
LoRA 微调(RF-DETR)
实验性 lora=True 会向 RF-DETR 主干注入 LoRA 适配器,用于 低显存微调。它需要可选的 peft 依赖 (pip install "libreyolo[lora]"),目前仅限 RF-DETR;其他系列会抛出明确错误,而不是忽略该标志。
1 model = LibreYOLO("LibreRFDETRs.pt") 2 results = model.train(data="data.yaml", epochs=50, lora=True)
实验日志记录器
v1.3.0 新增:传入 loggers= 可将指标流式发送到 TensorBoard、 MLflow 或 Weights & Biases。接受名称("tensorboard"、 "mlflow"、"wandb")、已配置的日志记录器 实例,或两者混合的可迭代对象。每个后端都是可选 extra (libreyolo[tensorboard]、[mlflow]、 [wandb])。
1 from libreyolo import LibreYOLO 2 from libreyolo.training.loggers import MLflowLogger 3 4 model = LibreYOLO("LibreYOLO9c.pt") 5 6 # By name 7 model.train(data="coco128.yaml", loggers="tensorboard") 8 9 # Mix configured instances and names 10 model.train( 11 data="coco128.yaml", 12 loggers=[MLflowLogger(experiment_name="my-exp"), "tensorboard"], 13 )
日志记录器仅为 Python API 功能。命令行没有相应标志;其余新增的训练参数 (--task、--cache、 --lora、--freeze、 --save-plots)在命令行中可用。
训练结果字典
1 { 2 "final_loss": 2.31, 3 "best_mAP50": 0.682, 4 "best_mAP50_95": 0.451, 5 "best_epoch": 87, 6 "save_dir": "runs/train/yolo9_exp", 7 "best_checkpoint": "runs/train/yolo9_exp/weights/best.pt", 8 "last_checkpoint": "runs/train/yolo9_exp/weights/last.pt", 9 }
恢复训练
1 # Load the checkpoint with the factory, then resume 2 model = LibreYOLO("runs/train/yolo9_exp/weights/last.pt") 3 results = model.train(data="coco128.yaml", resume=True)
自定义数据集 YAML 格式
1 path: /path/to/dataset 2 train: images/train 3 val: images/val 4 test: images/test # optional 5 6 nc: 3 7 names: ["cat", "dog", "bird"]
其他训练路径
其他系列也有训练器钩子,但在 v1.3.0 中它们不是推荐路径。新工作请保持使用 YOLO9 检测或 RF-DETR 检测/分割;仅出于兼容性、基准复现或针对性研究使用实验性训练器。PicoDet、RTMDet 和 EC 训练需要显式的 allow_experimental=True 确认。
从 YAML 配置训练
每个 model.train(...) 都接受 cfg="train.yaml",以从文件加载全部参数。显式 kwargs 仍优先于 yaml 中的值,因此你可以用 yaml 作为基线,并按运行覆盖单个字段。
1 model = LibreYOLO("LibreYOLO9c.pt") 2 results = model.train(cfg="configs/yolo9_finetune.yaml") 3 # Override individual fields: 4 # results = model.train(cfg="configs/yolo9_finetune.yaml", epochs=50)
梯度累积
传入 nbs(名义批大小)以启用梯度累积。训练器每 nbs / batch 次前向后步进一次优化器,从而让你在较小硬件上以配方的参考批大小进行训练。
1 # Effective batch 64 on a single GPU that only fits batch=8 2 model.train(data="coco128.yaml", batch=8, nbs=64)
分布式训练(DDP,实验性)
YOLO9 和 RF-DETR 通过 PyTorch DistributedDataParallel 支持多 GPU 训练,但多 GPU 不在 v1.3.0 经过充分测试的范围内。使用 torchrun 启动训练脚本:
1 # 4-GPU node 2 torchrun --nproc_per_node=4 train_yolo9.py 3 4 # Multi-node - see PyTorch's torchrun docs for --nnodes / --rdzv_endpoint
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreYOLO9c.pt") 4 # Pass device="" (auto-detect) and let torchrun set the rank 5 model.train(data="coco128.yaml", epochs=300, batch=16)
验证
在验证集上运行 COCO 标准评估。经过充分测试的验证路径是单 GPU 的 YOLO9 检测、RF-DETR 检测和 RF-DETR 分割。
1 results = model.val( 2 data="coco128.yaml", # dataset config 3 batch=16, 4 imgsz=640, 5 conf=0.001, # low conf for mAP calculation 6 iou=0.6, # NMS IoU threshold 7 split="val", # "val", "test", or "train" 8 save_json=False, # save predictions as COCO JSON 9 verbose=True, # print per-class metrics 10 plots=True, # save validation plots (metrics, per-class AP, confusion matrix); alias for save_plots 11 ) 12 13 print(f"mAP50: {results['metrics/mAP50']:.3f}") 14 print(f"mAP50-95: {results['metrics/mAP50-95']:.3f}")
验证结果字典
默认情况下,LibreYOLO 使用 COCO 评估,返回精确率、召回率、AP/AR 指标以及每张图像的耗时:
1 { 2 "metrics/mAP50-95": 0.489, # COCO primary metric (AP@[.5:.95]) 3 "metrics/mAP50": 0.721, # AP@0.5 (PASCAL VOC style) 4 "metrics/mAP75": 0.534, # AP@0.75 (strict) 5 "metrics/precision": 0.68, 6 "metrics/recall": 0.61, 7 "metrics/precision(B)": 0.68, # bbox aliases 8 "metrics/recall(B)": 0.61, 9 "metrics/mAP50(B)": 0.721, 10 "metrics/mAP50-95(B)": 0.489, 11 "metrics/mAP_small": 0.291, 12 "metrics/mAP_medium": 0.532, 13 "metrics/mAP_large": 0.648, 14 "metrics/AR1": 0.362, # Average Recall (max 1 det) 15 "metrics/AR10": 0.571, 16 "metrics/AR100": 0.601, 17 "metrics/AR_small": 0.387, 18 "metrics/AR_medium": 0.641, 19 "metrics/AR_large": 0.739, 20 "speed/preprocess_ms": 1.2, 21 "speed/inference_ms": 6.8, 22 "speed/postprocess_ms": 0.9, 23 "speed/total_ms": 8.9, 24 "speed/total_s": 12.3, 25 "speed/images_seen": 1382, 26 }
分割验证返回带 (M) 后缀的掩码指标,以及带 (B) 后缀的 框指标;OBB 验证会添加 (OBB) 指标。姿态验证通过 PoseValidator 返回 COCO 关键点指标。v1.3.0 新增了用于分类(top-1 / top-5)、 语义(mIoU / 像素精度)、点定位和深度(零样本)的验证器。传入 plots=True(或命令行的 --save-plots)可将 指标图、各类 AP、混淆矩阵和样本图写入运行目录。
导出
将 PyTorch 模型导出为 ONNX、TorchScript、TensorRT、OpenVINO、NCNN、CoreML 或(v1.3.0 新增)TFLite 以进行部署。TensorRT 现已覆盖所有模型系列。经过充分测试的导出路径仍是单 GPU 的 YOLO9 检测、RF-DETR 检测和 RF-DETR 分割。
快速导出
1 # ONNX (default) 2 model.export() 3 4 # TorchScript 5 model.export(format="torchscript") 6 7 # TensorRT (requires NVIDIA GPU + TensorRT) 8 model.export(format="tensorrt") 9 10 # OpenVINO (optimized for Intel hardware) 11 model.export(format="openvino") 12 13 # NCNN (via PNNX) 14 model.export(format="ncnn") 15 16 # CoreML (.mlpackage, macOS runtime) 17 model.export(format="coreml") 18 19 # TFLite (RF-DETR detect/seg/pose + YOLO9 detect; experimental, needs Python 3.12+) 20 model.export(format="tflite")
全部导出参数
1 path = model.export( 2 format="onnx", # "onnx", "torchscript", "tensorrt", "openvino", "ncnn", "coreml", or "tflite" 3 output_path="model.onnx", # output file (auto-generated if None) 4 imgsz=640, # input resolution (default: model's native); also accepts (h, w) for rectangular 5 opset=None, # ONNX opset (auto: 13, or 17 for wrappers that need it) 6 simplify=True, # run onnxsim graph simplification 7 dynamic=True, # enable dynamic batch axis (ONNX); TFLite requires static shapes 8 half=False, # export in FP16 9 batch=1, # batch size for static graph 10 device=None, # device to trace on (default: model's current device) 11 int8=False, # INT8 quantization: TensorRT, OpenVINO, or ONNX (YOLO9 detection only) 12 data=None, # calibration dataset for INT8 13 fraction=1.0, # fraction of calibration data to use 14 allow_download_scripts=False, # allow data.yaml download hooks during calibration 15 workspace=4.0, # TensorRT workspace size (GB) 16 min_batch=1, # TensorRT dynamic profile minimum batch 17 opt_batch=1, # TensorRT dynamic profile optimal batch 18 max_batch=8, # TensorRT dynamic profile maximum batch 19 hardware_compatibility="none", # TensorRT compatibility mode 20 gpu_device=0, # GPU device index for TensorRT 21 trt_config=None, # optional TensorRT YAML config path 22 compute_units="all", # CoreML routing: all, cpu_only, cpu_and_gpu, cpu_and_ne 23 nms=False, # embed NMS in the graph (ONNX YOLO9 detection, or CoreML) 24 iou=0.45, # embedded-NMS IoU threshold 25 conf=0.25, # embedded-NMS confidence threshold 26 max_det=300, # embedded-NMS max detections (ONNX only) 27 verbose=False, # verbose logging 28 )
OpenVINO INT8 导出额外需要 nncf。NCNN 导出会写出一个目录,包含 model.ncnn.param、model.ncnn.bin 和 metadata.yaml。CoreML 导出会写出 .mlpackage 包,需要 coremltools,且不支持 INT8。
ONNX 内嵌 NMS(YOLO9 检测)
v1.3.0 新增:传入 nms=True 可将 NMS 烘焙进导出的 ONNX 图,使模型直接输出最终框。目前仅限 detect 任务上的 yolo9 系列(其他系列/任务会报错)。它会强制固定的 batch-1 图 (dynamic=False),并在 ONNX 元数据中记录 nms / nms_conf / nms_iou / max_det。
1 model = LibreYOLO("LibreYOLO9c.pt") 2 model.export(format="onnx", nms=True, conf=0.25, iou=0.45, max_det=300)
int8=True 现在也支持 ONNX(除 TensorRT 和 OpenVINO 之外), 同样仅限 YOLO9 检测;它需要一个用于标定的 data= 数据集。
TFLite (LiteRT) 导出
实验性 v1.3.0 新增了基于 onnx2tf 的 TFLite 导出路径。TFLite 是 Google LiteRT 运行时的格式(TensorFlow Lite 于 2024 年更名为 LiteRT,.tflite 文件格式不变)。它已在 RF-DETR detect / segment / pose 和 YOLO9 detect 上验证。它需要 Python 3.12+( onnx2tf 2.4.x wheels 不面向更老的 Python),以及 可选 extra libreyolo[tflite](onnx2tf>=2.4.3、onnx-graphsurgeon、onnx-simplifier)。导出 为 FP32 且仅静态形状(暂不支持 half、 int8 或 dynamic)。
1 pip install "libreyolo[tflite]" # Python 3.12+
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreRFDETRs-seg.pt") 4 model.export(format="tflite") # writes a .tflite file
对于 RF-DETR,导出器会将每个 GridSample 节点改写为 TFLite 安全的双线性 子图,因为 onnx2tf 的默认 lowering 在数值上有问题。在 v1.3.0 中,由于 onnx2tf>=2.4.3 已在上游提供 RF-DETR 修复,旧的针对 onnx2tf 的运行时 monkeypatch 已被移除;只保留静态 ONNX 图改写。
没有 TFLite 运行时后端。 LibreYOLO 无法加载或运行 .tflite 文件;该格式仅用于 导出。请在目标设备上使用 Google 的 LiteRT 运行时 (pip install ai-edge-litert)运行导出的模型。请避免使用 旧的 tflite-runtime 包,它已不再更新。
ONNX 元数据
导出的 ONNX 文件包含内嵌元数据:
| 键 | 示例值 |
|---|---|
libreyolo_version | "1.3.0" |
model_family | "yolox" |
model_size | "s" |
nb_classes | "80" |
names | '{"0": "person", "1": "bicycle", ...}' |
imgsz | "640" |
dynamic | "True" |
half | "False" |
使用 LibreYOLO("model.onnx") 加载导出文件时,会自动读回该元数据。
TorchScript 推理
通过相同的运行时后端预测 API 运行导出的 .torchscript 模型。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model.torchscript") 4 5 result = model("image.jpg", conf=0.25, iou=0.45, save=True) 6 print(result.boxes.xyxy)
ONNX 推理
使用 ONNX Runtime 而非 PyTorch 运行推理。适用于没有 PyTorch 的部署环境。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model.onnx") 4 5 result = model("image.jpg", conf=0.25, iou=0.45, save=True) 6 print(result.boxes.xyxy)
自动元数据
如果 ONNX 文件由 LibreYOLO 导出,类别名称和类别数会自动从内嵌元数据中读取:
1 # Export with metadata 2 model.export(format="onnx", output_path="model.onnx") 3 4 # Load - names and nb_classes auto-populated 5 onnx_model = LibreYOLO("model.onnx") 6 print(onnx_model.names) # {0: "person", 1: "bicycle", ...} 7 print(onnx_model.nb_classes) # 80
对于没有元数据的 ONNX 文件(例如由其他工具导出),请手动指定 nb_classes:
1 model = LibreYOLO("external_model.onnx", nb_classes=20)
设备选择
1 # Auto-detect (CUDA if available, else CPU) 2 model = LibreYOLO("model.onnx", device="auto") 3 4 # Force CPU 5 model = LibreYOLO("model.onnx", device="cpu") 6 7 # Force CUDA 8 model = LibreYOLO("model.onnx", device="cuda")
预测参数
通过 LibreYOLO() 加载的运行时产物支持共享的运行时预测 API:
1 result = model( 2 "image.jpg", 3 conf=0.25, 4 iou=0.45, 5 imgsz=640, 6 classes=[0, 2], 7 max_det=300, 8 save=True, 9 output_path="output/annotated.jpg", # final file path when save=True 10 color_format="auto", 11 )
运行时后端不暴露仅 PyTorch 的选项,如 tiling、overlap_ratio 或 output_file_format。
运行时后端在保存方面与 PyTorch 包装器略有不同:如果你设置 output_path,请传入最终文件路径,而非目录。如果省略它,当前后端默认保存在 runs/detections/ 下。
TensorRT 推理
使用 TensorRT 在 NVIDIA GPU 上获得最大吞吐。需要 CUDA 以及 TensorRT 的 Python 绑定。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model.engine") 4 5 result = model("image.jpg", conf=0.25, iou=0.45, save=True) 6 print(result.boxes.xyxy)
通过 LibreYOLO() 加载的 TensorRT 产物支持与 ONNX 和 OpenVINO 相同的核心运行时预测 API,包括 save=True 时同样仅接受文件路径的 output_path 行为。
OpenVINO 推理
使用 OpenVINO 运行推理,针对 Intel CPU、GPU 和 VPU 优化。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model_openvino/") 4 5 result = model("image.jpg", conf=0.25, iou=0.45, save=True) 6 print(result.boxes.xyxy)
通过 LibreYOLO() 加载的 OpenVINO 目录会在存在时读取 metadata.yaml,并支持相同的核心运行时预测 API。
NCNN 推理
使用 NCNN 运行推理,用于在 CPU 或支持 Vulkan 的 GPU 目标上进行轻量部署。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model_ncnn/") 4 5 result = model("image.jpg", conf=0.25, iou=0.45, save=True) 6 print(result.boxes.xyxy)
NCNN 导出目录包含 model.ncnn.param、model.ncnn.bin,通常还有 metadata.yaml。
CoreML 推理
在 macOS 上通过 CoreML 运行导出的 .mlpackage。CoreML 使用 compute_units 而非 PyTorch 设备字符串来路由执行。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model.mlpackage", compute_units="all") 4 5 result = model("image.jpg", conf=0.25, iou=0.45, save=True) 6 print(result.boxes.xyxy)
支持的 compute_units 取值为 all、cpu_only、cpu_and_gpu 和 cpu_and_ne。
命令行(CLI)
安装 LibreYOLO 会在你的 PATH 中注册一个 libreyolo 命令(入口点在 pyproject.toml 中)。命令行镜像 Python API,并遵循 Ultralytics 风格的 key=value 语法。
子命令
| 命令 | 用途 |
|---|---|
predict | 对图像、目录或视频运行推理 |
train | 在数据集上训练模型 |
val | 在数据集上评估模型 |
export | 导出为 ONNX / TorchScript / TensorRT / OpenVINO / NCNN / CoreML / TFLite |
ui | 启动本地拖放 / 粘贴的浏览器推理界面 |
doctor | 运行训练前的数据集健康检查(YOLO 检测格式) |
checks | 打印 Python、torch、CUDA、GPU 和可选包信息 |
models | 列出已注册的模型系列和命令行简称 |
formats | 列出支持的导出格式 |
cfg | 打印默认的训练配置 YAML |
info | 加载模型并打印解析出的系列、尺寸、任务、设备和类别 |
metadata | 检查 .pt 文件中的原始检查点元数据 |
version | 打印 LibreYOLO + Python + torch 版本 |
模型名称简称
命令行接受简称(yolo9-c),它们会解析为权重文件名(LibreYOLO9c.pt)- 可通过 libreyolo models 查看。你也可以传入任意明确的检查点路径。
常用选项
| 命令 | 重要选项 |
|---|---|
predict | conf, iou, imgsz, classes, max_det, half, batch, tiling, overlap_ratio, output_file_format, project, name, exist_ok, face_detector |
train | epochs, batch, imgsz, lr0, optimizer, scheduler, workers, seed, resume, amp, task, cache, lora, freeze, save_plots, allow_download_scripts, dry_run |
val | split, batch, imgsz, conf, iou, max_det, half, save_plots, data_dir, use_coco_eval, project, name, exist_ok, save_json, allow_download_scripts |
export | format, imgsz, batch, half, int8, dynamic, simplify, nms, conf, iou, max_det, opset, data, fraction, device, allow_download_scripts, verbose |
预测
1 # Flagship: YOLO9 2 libreyolo predict model=yolo9-c source=image.jpg conf=0.25 save=true 3 4 # Flagship: RF-DETR 5 libreyolo predict model=rfdetr-s source=image.jpg save=true 6 7 # Video - saved under runs/detect/predict*/ 8 libreyolo predict model=yolo9-c source=clip.mp4 save=true 9 10 # Tiled inference for very large images 11 libreyolo predict model=yolo9-c source=aerial.jpg tiling=true save=true 12 13 # Gaze (requires a face detector) 14 libreyolo predict model=LibreL2CSr50.pt source=portrait.jpg \ 15 --face-detector path/to/face.pt save=true
训练
1 libreyolo train model=yolo9-c data=coco128.yaml epochs=300 batch=16 device=0 2 3 # Dry-run prints the resolved config without launching training 4 libreyolo train model=yolo9-c data=coco128.yaml --dry-run
验证
1 libreyolo val model=runs/train/exp/weights/best.pt data=coco128.yaml split=val
导出
1 libreyolo export model=runs/train/exp/weights/best.pt format=onnx dynamic=true 2 libreyolo export model=best.pt format=tensorrt half=true 3 libreyolo export model=best.pt format=openvino int8=true data=coco128.yaml 4 libreyolo export model=best.pt format=coreml
导出内嵌 NMS 与矩形尺寸
1 # Embed NMS into an ONNX YOLO9 detection graph 2 libreyolo export model=yolo9-c format=onnx nms=true conf=0.25 iou=0.45 max_det=300 3 4 # Rectangular export size (imgsz accepts a single value or two comma-separated dims) 5 libreyolo export model=yolo9-c format=onnx imgsz=640,480 6 7 # TFLite (Python 3.12+, libreyolo[tflite]) 8 libreyolo export model=rfdetr-s format=tflite
本地推理界面
libreyolo ui 提供一个本地浏览器页面,你可以在其中拖放、粘贴 或选择图像、选择模型并查看结果。它默认绑定 127.0.0.1:8000,若端口被占用会自动递增。
1 libreyolo ui # opens http://127.0.0.1:8000 2 libreyolo ui --port 9000 --no-browser --device 0
数据集健康检查
libreyolo doctor 对 YOLO 检测格式的数据集运行训练前检查, 发现错误时以非零退出 (--strict 还会在出现警告时失败),因此可用于门控 CI。
1 libreyolo doctor coco8.yaml 2 libreyolo doctor --data coco8.yaml --strict --json 3 libreyolo doctor coco8.yaml --fast --only labels # skip image decoding, run one check family
机器可读输出
每个命令都接受 --json(结构化的标准输出,便于管道传入脚本或智能体)和 --quiet(抑制 stderr 进度行)。核心的 predict、train、val 和 export 命令还接受 --help-json,可将其参数 schema 以 JSON 形式输出。
1 libreyolo predict model=yolo9-c source=img.jpg --json | jq . 2 3 libreyolo train --help-json > train_schema.json
API 参考
LibreYOLO(工厂)
1 LibreYOLO( 2 model_path: str, 3 *, 4 device: str = "auto", 5 task: str | None = None, # override only when a custom artifact is ambiguous 6 nb_classes: int | None = None, # mainly for external exported artifacts 7 compute_units: str = "all", # CoreML only: all, cpu_only, cpu_and_gpu, cpu_and_ne 8 ) -> model wrapper or runtime backend
优先使用官方检查点文件名和导出产物路径,然后让工厂解析细节。它可处理 PyTorch 检查点、.onnx、.torchscript、.engine、.tensorrt、.mlpackage、包含 model.xml 的 OpenVINO 目录,以及包含 model.ncnn.param 和 model.ncnn.bin 的 NCNN 目录。task 参数用于有歧义的自定义产物;否则解析来自检查点元数据、文件名后缀和系列默认值。
预测(PyTorch 模型包装器)
1 model( 2 source, # image input (see supported formats) 3 *, 4 conf: float = 0.25, 5 iou: float = 0.45, 6 imgsz: int = None, 7 device: str = "auto", 8 classes: list[int] = None, 9 max_det: int = 300, 10 augment: bool = False, 11 save: bool = False, 12 batch: int = 1, 13 stream: bool = False, 14 vid_stride: int = 1, 15 show: bool = False, 16 output_path: str = None, 17 color_format: str = "auto", 18 tiling: bool = False, 19 overlap_ratio: float = 0.2, 20 output_file_format: str = None, 21 ) -> Results | list[Results] | Generator[Results, None, None]
预测(运行时后端)
1 backend( 2 source, 3 *, 4 conf: float = 0.25, 5 iou: float = 0.45, 6 imgsz: int = None, 7 classes: list[int] = None, 8 max_det: int = 300, 9 save: bool = False, 10 batch: int = 1, 11 output_path: str = None, # final file path when save=True 12 color_format: str = "auto", 13 ) -> Results | list[Results]
如果运行时后端省略了 output_path,当前默认保存位置是 runs/detections/。
Results
1 result = Results( 2 boxes: Boxes | None, 3 orig_shape: tuple[int, int], # (height, width) 4 path: str | None, 5 names: dict[int, str], 6 masks: Masks | None = None, 7 keypoints: Keypoints | None = None, 8 probs: Probs | None = None, 9 obb: OBB | None = None, 10 gaze: Gaze | None = None, 11 speed: dict[str, float] | None = None, 12 track_id = None, 13 frame_idx: int | None = None, 14 ) 15 16 len(result) # number of detections 17 result.cpu() # copy with tensors on CPU 18 result.cuda() # copy with tensors on CUDA 19 result.numpy() # copy with numpy arrays 20 result.summary() # list[dict] with boxes, masks, gaze, and track_id when present 21 result.to_json() # JSON string from summary()
Boxes
1 boxes = Boxes(boxes, conf, cls) 2 3 boxes.xyxy # (N, 4) tensor - x1, y1, x2, y2 4 boxes.xywh # (N, 4) tensor - cx, cy, w, h 5 boxes.conf # (N,) tensor - confidence scores 6 boxes.cls # (N,) tensor - class IDs 7 boxes.id # (N,) track IDs when tracking, else None 8 boxes.is_track # True when track IDs are attached 9 boxes.data # (N, 6) [xyxy, conf, cls], or (N, 7) with track IDs 10 11 len(boxes) # number of boxes 12 boxes.cpu() # copy on CPU 13 boxes.numpy() # copy as numpy arrays
任务载荷
1 result.masks.data # segmentation masks, (N, H, W) 2 result.masks.xy # list of mask contours in pixel coordinates 3 result.masks.xyn # normalized mask contours 4 5 result.keypoints.xy # pose keypoint coordinates 6 result.keypoints.xyn # normalized keypoint coordinates 7 result.keypoints.conf # keypoint confidence when present 8 9 result.gaze.data # (N, 2): pitch, yaw in radians 10 result.gaze.pitch_deg # pitch in degrees 11 result.gaze.yaw_deg # yaw in degrees 12 result.gaze.direction_3d # approximate 3D direction vectors
model.export()
1 model.export( 2 format: str = "onnx", # "onnx", "torchscript", "tensorrt", "openvino", "ncnn", or "coreml" 3 *, 4 output_path: str | None = None, 5 imgsz: int | None = None, 6 opset: int | None = None, # auto: 13, or 17 for wrappers that need it 7 simplify: bool = True, 8 dynamic: bool = True, 9 half: bool = False, 10 batch: int = 1, 11 device: str | None = None, 12 int8: bool = False, 13 data: str | None = None, # calibration data for INT8 14 fraction: float = 1.0, # fraction of calibration data 15 allow_download_scripts: bool = False, 16 workspace: float = 4.0, # TensorRT workspace (GB) 17 min_batch: int = 1, # TensorRT dynamic profile minimum batch 18 opt_batch: int = 1, # TensorRT dynamic profile optimal batch 19 max_batch: int = 8, # TensorRT dynamic profile maximum batch 20 hardware_compatibility: str = "none", 21 gpu_device: int = 0, 22 trt_config = None, # optional TensorRT YAML config path 23 compute_units: str = "all", # CoreML only 24 nms: bool = False, # CoreML embedded NMS where supported 25 iou: float = 0.45, # CoreML embedded NMS IoU threshold 26 conf: float = 0.25, # CoreML embedded NMS confidence threshold 27 verbose: bool = False, 28 ) -> str # path to exported file or directory
model.val()
1 model.val( 2 data: str = None, # path to data.yaml 3 batch: int = 16, 4 imgsz: int = None, 5 conf: float = 0.001, 6 iou: float = 0.6, 7 workers: int = 4, 8 allow_download_scripts: bool = False, 9 device: str = None, 10 split: str = "val", # "val", "test", or "train" 11 augment: bool = False, 12 save_json: bool = False, 13 verbose: bool = True, 14 ) -> dict
返回(COCO 评估,默认):
1 { 2 "metrics/mAP50-95": float, # COCO primary metric 3 "metrics/mAP50": float, 4 "metrics/mAP75": float, 5 "metrics/mAP_small": float, 6 "metrics/mAP_medium": float, 7 "metrics/mAP_large": float, 8 "metrics/AR1": float, 9 "metrics/AR10": float, 10 "metrics/AR100": float, 11 "metrics/AR_small": float, 12 "metrics/AR_medium": float, 13 "metrics/AR_large": float, 14 }
model.train()(YOLO9)
1 model.train( 2 data: str, # path to data.yaml (required) 3 *, 4 epochs: int = 300, 5 batch: int = 16, 6 imgsz: int = 640, 7 lr0: float = 0.01, 8 optimizer: str = "SGD", 9 device: str = "", 10 workers: int = 8, 11 seed: int = 0, 12 project: str = "runs/train", 13 name: str = "yolo9_exp", 14 exist_ok: bool = False, 15 resume: bool = False, 16 amp: bool = True, 17 patience: int = 50, 18 allow_download_scripts: bool = False, 19 callbacks = None, 20 ) -> dict
返回标准的 LibreYOLO 训练字典,包含 final_loss、best_mAP50、best_mAP50_95、best_epoch、save_dir、best_checkpoint 和 last_checkpoint。
model.train()(RF-DETR)
1 model.train( 2 data: str, # path to data.yaml 3 epochs: int = 100, 4 batch_size: int = 4, 5 lr: float = 1e-4, 6 output_dir: str = "runs/train", 7 resume: str = None, 8 **kwargs, # additional RF-DETR training args 9 ) -> dict
YOLO-NAS、D-FINE、DEIM、DEIMv2、EC、PicoDet、RT-DETRv2/v4 和 RTMDet 还有额外的实验性训练器,外加新增的分类(MobileNetV4、ConvNeXt、EfficientNetV2、DINOv2)、语义分割(DINOv2)和点定位(FOMO)系列。它们遵循相同的 model.train(data="...yaml", ...) 形式,但默认值和实验性门控因系列而异。
运行时产物加载
通过 LibreYOLO() 加载导出产物,方式与加载 PyTorch 检查点相同。工厂会根据路径选择 ONNX Runtime、TorchScript、TensorRT、OpenVINO、NCNN 或 CoreML:
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("model.onnx") 4 model = LibreYOLO("model.torchscript") 5 model = LibreYOLO("model.engine") 6 model = LibreYOLO("model_openvino/") 7 model = LibreYOLO("model_ncnn/") 8 model = LibreYOLO("model.mlpackage", compute_units="all")
高级集成可以触及更底层的运行时模块,但普通应用代码应保持使用工厂路径。
ValidationConfig
1 from libreyolo import ValidationConfig 2 3 config = ValidationConfig( 4 data="coco128.yaml", 5 data_dir=None, # override dataset root directory 6 split="val", # "val", "test", or "train" 7 batch_size=16, 8 imgsz=640, 9 conf_thres=0.001, 10 iou_thres=0.6, 11 max_det=300, 12 iou_thresholds=( # mAP IoU sweep 13 0.50, 0.55, 0.60, 0.65, 0.70, 0.75, 0.80, 0.85, 0.90, 0.95, 14 ), 15 device="auto", 16 save_dir=None, 17 save_json=False, 18 verbose=True, 19 num_workers=4, 20 half=False, 21 augment=False, # test-time augmentation (TTA) 22 allow_download_scripts=False, 23 # Pose-only fields (PoseValidator) 24 keypoints_json=None, 25 images_dir=None, 26 oks_sigmas=None, 27 ) 28 29 # Load/save YAML 30 config = ValidationConfig.from_yaml("config.yaml") 31 config.to_yaml("config.yaml")
架构指南
本节面向希望了解代码库内部机制的贡献者。
基类设计
PyTorch 模型系列继承自 libreyolo/models/base/model.py 中的 BaseModel。子类实现以下抽象方法:
| 方法 | 用途 |
|---|---|
_init_model() | 构建并返回 nn.Module |
_get_available_layers() | 返回层名到模块的映射 |
_get_preprocess_numpy() | 返回用于导出 / 标定的 NumPy 预处理器 |
_preprocess() | 图像到张量的转换 |
_forward() | 模型前向传播 |
_postprocess() | 原始输出到检测字典 |
BaseModel 提供共享的包装行为:预测、导出、验证、尺寸/名称元数据和训练辅助。实际的单图、批量和分块推理流程位于 libreyolo/models/base/inference.py,而部署运行时位于 libreyolo/backends/ 下。
包结构
1 libreyolo/ 2 __init__.py # Public API exports + deprecated-alias resolver 3 tasks.py # Task types, suffix conventions, resolution rules 4 assets/parkour.jpg # SAMPLE_IMAGE 5 models/ 6 __init__.py # LibreYOLO() factory + model registry bootstrap 7 base/ 8 model.py # BaseModel - shared wrapper behaviour 9 inference.py # Shared prediction pipeline (image/dir/video/tiled) 10 yolox/ # LibreYOLOX (detect) 11 yolo9/ # LibreYOLO9 (detect) 12 yolo9_e2e/ # LibreYOLO9E2E (detect) 13 yolonas/ # LibreYOLONAS (detect, pose) 14 dfine/ # LibreDFINE (detect) 15 deim/ # LibreDEIM (detect) 16 deimv2/ # LibreDEIMv2 (detect) 17 rtdetr/ # LibreRTDETR (detect) 18 rtdetrv2/ # LibreRTDETRv2 (detect) 19 rtdetrv4/ # LibreRTDETRv4 (detect) 20 rfdetr/ # LibreRFDETR (detect, segment, pose, obb) - lazy-loaded 21 ec/ # LibreEC / EdgeCrafter (detect, pose, segment) 22 picodet/ # LibrePICODET (detect) 23 rtmdet/ # LibreRTMDet (detect) 24 dinov2/ # LibreDINOv2 (semantic, classify) - lazy-loaded 25 mobilenetv4/ # LibreMobileNetV4 (classify) 26 convnext/ # LibreConvNeXt (classify) 27 efficientnetv2/ # LibreEfficientNetV2 (classify) 28 depth_anything/ # LibreDepthAnythingV2 (depth) 29 fomo/ # LibreFOMO (point) 30 l2cs/ # LibreL2CS (gaze, inference-only) 31 backends/ 32 base.py 33 onnx.py # ONNX Runtime loader 34 torchscript.py # TorchScript loader 35 tensorrt.py # TensorRT loader 36 openvino.py # OpenVINO loader 37 ncnn.py # NCNN loader 38 coreml.py # CoreML loader 39 export/ 40 exporter.py # BaseExporter and format registry 41 onnx.py / torchscript.py / tensorrt.py / openvino.py / ncnn.py / coreml.py 42 config.py / calibration.py 43 training/ 44 trainer.py # Shared trainer scaffolding 45 config.py # TrainConfig dataclass (single source of truth) 46 augment.py / callbacks.py / distributed.py / ema.py / scheduler.py 47 artifacts.py / train_config.yaml 48 # Per-family trainers live in models/<family>/trainer.py 49 validation/ 50 config.py # ValidationConfig 51 base.py / preprocessors.py 52 detection_validator.py # DetectionValidator, SegmentationValidator 53 pose_validator.py # PoseValidator 54 coco_evaluator.py # COCOEvaluator 55 tracking/ 56 tracker.py # ByteTracker 57 config.py # TrackConfig 58 kalman_filter.py / matching.py / strack.py 59 cli/ 60 __init__.py # libreyolo entrypoint (Typer app) 61 commands/ # predict / train / val / export / special 62 aliases.py / config.py / parsing.py / output.py / errors.py 63 utils/ 64 results.py # Results, Boxes, Masks, Keypoints, Probs, OBB, Gaze 65 image_loader.py # Unified image loading 66 video.py # VideoSource, VideoWriter, video inference loop 67 general.py # Path helpers, NMS, tiling utilities 68 download.py / drawing.py / logging.py / predict_args.py 69 serialization.py / box_ops.py 70 data/ 71 dataset.py / pose_dataset.py / utils.py / yolo_coco_api.py 72 config/ 73 datasets/ # Built-in dataset YAML configs (coco8, coco128, coco5000, coco, etc.) 74 export/ # TensorRT default YAML
添加新的模型系列
- 1创建
libreyolo/models/newmodel/model.py,其中包含一个继承BaseModel的类 - 2按需设置
FAMILY、FILENAME_PREFIX、INPUT_SIZES、SUPPORTED_TASKS和DEFAULT_TASK - 3实现注册表钩子,如
can_load()、detect_size()、detect_nb_classes()和detect_size_from_filename() - 4实现该系列所需的模型初始化、预处理、前向、后处理、训练和验证钩子
- 5在
libreyolo/models/newmodel/下创建配套的网络和工具 - 6将导入添加到
libreyolo/models/__init__.py;子类注册会在导入运行时发生 - 7从
libreyolo/__init__.py导出该类 - 8(可选)如果验证预处理与标准路径不同,覆盖
val_preprocessor_class
导出架构
用户代码应通过 model.export(...) 导出。在内部,libreyolo/export/exporter.py 中的 BaseExporter 拥有格式注册表,具体的导出器通过子类注册来注册自身。
1 from libreyolo import LibreYOLO 2 3 model = LibreYOLO("LibreYOLO9c.pt") 4 model.export(format="onnx")
要添加新的导出格式,请实现一个具有唯一 format_name 的新 BaseExporter 子类,并从 libreyolo/export/exporter.py 导入它,以填充注册表。
数据集格式
每个任务都通过一个 data.yaml 加载。检测、实例分割和 OBB 接受两种可互换的标签格式——YOLO TXT 或原生 COCO JSON——加载器会根据配置自动选择。姿态、语义分割、深度和分类各自增加一个小型格式。下表将每个任务映射到其布局。
按任务划分的格式
| 任务 | 数据布局 | 标签 |
|---|---|---|
| 检测 | data.yaml + labels/*.txt,或 COCO JSON | 每行一个框 |
| 实例分割 | data.yaml + 多边形 .txt,或 COCO JSON | 每行一个多边形(TXT)/ 多边形 + RLE(COCO) |
| OBB | data.yaml + 旋转框 .txt,或 COCO JSON | 每行一个旋转框 |
| 姿态 | data.yaml + .txt + kpt_shape/flip_idx | 每行一个框 + 关键点 |
| 语义分割 | data.yaml + masks_dir/ PNG | 逐像素类别 ID(255 = 忽略) |
| 深度 | data.yaml + depths_dir/ 深度图 | 逐像素深度(0 = 无效) |
| 分类 | ImageFolder(train/<class>/) | 文件夹名 = 类别 |
data.yaml 结构
检测、分割、OBB 和姿态的通用契约。train/val/test 可以是目录、.txt 文件列表(每行一个图像路径)或路径列表。nc 可选——省略时从 names 推断。
1 path: /absolute/path/to/dataset # dataset root 2 train: images/train # dir, .txt file list, or list of paths 3 val: images/val 4 test: images/test # optional 5 6 nc: 80 # optional; inferred from names if absent 7 names: ["person", "bicycle", "car", "..."]
配置按以下顺序解析:明确路径、当前工作目录,然后是 libreyolo/config/datasets/ 下的内置配置。根目录默认在 ~/datasets 下;用 LIBREYOLO_DATASETS_DIR 覆盖。
YOLO TXT 标签
默认布局:每张图像在 labels/ 下有一个 .txt,镜像 images/ 树并使用相同的文件名主干。所有坐标都归一化到 [0, 1]。
1 dataset/ 2 images/train/img001.jpg 3 labels/train/img001.txt # same stem as the image
1 # Detection one box per line 2 <class_id> <cx> <cy> <w> <h> 3 4 # Segmentation one polygon per line (box derived from the vertices) 5 <class_id> <x1> <y1> <x2> <y2> ... <xn> <yn> 6 7 # Pose box, then K keypoints (needs kpt_shape / flip_idx below) 8 <class_id> <cx> <cy> <w> <h> <kx1> <ky1> <v1> ... <kxK> <kyK> <vK> 9 10 # OBB four rotated-box corners 11 <class_id> <x1> <y1> <x2> <y2> <x3> <y3> <x4> <y4>
1 kpt_shape: [17, 3] # K keypoints, 3 values each: x, y, visibility 2 flip_idx: [0, 2, 1, 4, 3, 6, 5, 8, 7, 10, 9, 12, 11, 14, 13, 16, 15]
原生 COCO JSON
检测、分割和 OBB 也可直接加载 COCO JSON——添加一个 annotations: 块,将每个划分映射到其 JSON 文件。此时 train/val 指向图像目录(而非 .txt 列表)。需要 pycocotools;类别名来自 JSON 的 categories,因此 nc/names 可选。
1 path: dataset 2 train: images/train # image directory 3 val: images/val 4 annotations: 5 train: annotations/train.json # COCO instances JSON 6 val: annotations/val.json
同一套切换逻辑服务于 YOLO9、RF-DETR、DEIM 和 D-FINE 训练,以及检测、OBB 和姿态验证器。磁盘上带有 annotations/instances_train2017.json 的 COCO 布局也会被自动识别,无需 annotations: 键。
选择哪种分割格式? 一个 YOLO 多边形行是每个实例一个环——无法表示带孔或分离(多部分)的掩码。COCO JSON 会保留实例的每个多边形并解码 RLE 掩码(含孔)。当实例带孔或存在不相连的部分时使用 COCO JSON;简单的单块掩码用任一格式均可。人群标注(iscrowd: 1)会被跳过。
语义分割掩码
为每张图像配一张单通道掩码,其像素值为类别 ID;255 标记被忽略的像素。masks_dir 会替换每个路径中的 images(默认 masks),掩码必须是无损(PNG)且与其图像同名主干。可选的 label_mapping 将源 ID 重映射为训练 ID(未映射的值变为忽略)。省略 masks_dir 则在加载时从 YOLO 多边形标签栅格化掩码,并追加一个 background 类。
1 dataset/ 2 images/train/scene001.jpg 3 masks/train/scene001.png # single-channel class IDs, 255 = ignore
1 path: /path/to/dataset 2 train: images/train 3 val: images/val 4 masks_dir: masks 5 nc: 3 6 names: ["road", "building", "vegetation"]
深度图
为每张图像配一张单通道深度图,放在 depths_dir 下(默认 depths)。16 位 PNG/TIF 会除以 depth_scale(默认 256.0);.npy 浮点文件按原样使用。零、负值和非有限像素为无效。可选的 depth_stem_suffix 和 *_mask 有效性掩码会被自动识别。深度在 v1.3.0 中仅支持验证。
1 path: /path/to/dataset 2 val: images/val 3 depths_dir: depths 4 depth_scale: 256.0 # 16-bit PNG encoding: value / 256 = depth
分类
分类使用 ImageNet 风格的 ImageFolder 树,而非 data.yaml——布局参见 分类。data= 接受数据集根目录、.zip URL 或已知名称。
内置数据集
配置位于 libreyolo/config/datasets/ 下。下载行为因配置而异:基于 URL 的数据集在首次使用时获取,基于脚本的数据集需要 allow_download_scripts=True,少数需要在本地放置。
| 配置 | 任务 | 下载 |
|---|---|---|
coco8 | 检测(8 张) | 自动 |
coco128 | 检测(128 张) | 自动 |
coco5000 | 检测 | 脚本 —— allow_download_scripts=True |
coco / coco-val-only | 检测(完整) | 脚本 —— allow_download_scripts=True |
coco8-pose / coco-pose | 姿态 | 脚本 —— allow_download_scripts=True |
cocostuff | 语义(182 类) | 手动 —— 本地放置 |
1 results = model.val(data="coco8.yaml") # auto-downloads 2 results = model.train(data="coco128.yaml", epochs=10) # auto-downloads 3 model.train(data="coco8-pose.yaml", allow_download_scripts=True) # script config