查看 Markdown

Grounding DINO

Grounding DINO 是一个开集(open-set)目标检测器,由 IDEA Research 开发,它拿图像去和一段自由文本提示词打分,而不是去对一个固定的类别列表。LibreYOLO 把它包装成开放词汇检测器层里一个仅支持预测的家族。

任务
detection
尺寸
t, b at 800 px
安装
pip install libreyolo
支持层级
兄弟层级,自 v 起。一个独立的产品面,有自己的工厂和契约。
上游
Grounding DINO,由 IDEA Research 发布,采用 Apache-2.0 许可。论文源码
许可
代码采用 MIT,权重采用 Apache-2.0。商用

安装

Grounding DINO 通过 LibreYOLO 的开放词汇检测器层加载,这一层需要 openvocab extra:

bash
pip install "libreyolo[openvocab]"

这个 extra 会装上 transformerstimm,也就是这一层调用的那两个 Hugging Face 库。

预测

Grounding DINO 不是 LibreYOLO 通过 LibreYOLO() 加载的检查点(checkpoint)。它 通过同级的 LibreOpenVocab 工厂加载,后者在首次使用时下载一份 Hugging Face 快照,并缓存在 weights/ 下。

Python
from libreyolo import LibreOpenVocab, SAMPLE_IMAGE model = LibreOpenVocab("grounding-dino-t")model.set_classes(["person", "dog", "skateboard"]) result = model.predict(SAMPLE_IMAGE, conf=0.25)for box in result.boxes:    print(box.cls, box.conf, box.xyxy)
文本阈值
from libreyolo import LibreOpenVocab, SAMPLE_IMAGE model = LibreOpenVocab("grounding-dino-b")model.set_classes(["remote control", "school bus"]) # conf 按检测框分数过滤,text_threshold 按解码出的短语的 token# 分数过滤,两者不设置时都默认为 0.25result = model.predict(SAMPLE_IMAGE, conf=0.25, text_threshold=0.3)print(result.names)

set_classes() 设定一个粘住不放的文本词汇表:再调用一次就能替换整个列表,不调用 则保留默认的 COCO-80 标签。Grounding DINO 从自己的文本输出里解码出自由形式的短 语,并自行把它们映射回那个词汇表,归一化之后的精确匹配优先,整 token 的匹配也会 被接受,而有歧义或匹配不上的短语会被丢弃,而不是靠猜,所以 school bus 绝不会 被映射成单独的 busschool。词汇表长到超出文本编码器的 token 上限时,会被 拆成若干个提示词,作为多次独立的前向传播运行,再合并回一组检测结果,并由 max_det 封顶。

iou 为了 API 兼容性会被接受,但会发出警告并且没有任何作用,因为这里没有任何东 西跑非极大值抑制。imgszaugment=True 会被直接拒绝:缩放由 transformers 的 processor 负责,而测试时增强不在这一层的范围内。对单张图像调用 predict() 返回的是一个 Results,不是列表;传入一个目录、一个图像列表,或者 对视频源用 stream=True,才会得到多个。这个家族没有 CLI 路径, libreyolo predict 只通过 LibreYOLO() 加载 .pt 检查点,所以 LibreOpenVocab 家族从 Python 里运行。数据源类型和流式处理见预测

变体

两个检查点,tbt 是这一层未指定尺寸时的默认尺寸。两者都通过 transformersGroundingDinoForObjectDetection 镜像官方的 IDEA Research 发布版,一次性下载到一份保留了上游文件的 LibreYOLO 托管 Hugging Face 快照里。 这个家族目前还没有发布任何精度或延迟数字。

训练、数据集验证和导出都不在这一层的范围内:train()val()export() 都 无条件抛出 NotImplementedError。这是一个围绕已发布检查点的、仅支持预测的封装。

检查点

这个家族已发布的全部权重文件。

文件输入(px)权重许可
Detection
LibreGroundingDINOt.pt800apache-2.0
LibreGroundingDINOb.pt800apache-2.0

上面的每个文件目前都在 LibreYOLO 组织中,并会在首次使用时下载。

许可证

请检查你所下载的具体权重在 Hugging Face 仓库中的许可。LibreYOLO 组织里的每个检查点都附有许可,同一家族内也不一定相同。该仓库是权威来源;以下摘要说明本页上次验证时适用的情况。

这里只说明涉及的许可证,不构成法律意见。如果答案对商用很重要,请自行阅读许可证并咨询法律顾问。

原始工作
Grounding DINO, IDEA Research
上游许可
Apache-2.0
LibreYOLO 代码
MIT
权重
采用 Apache-2.0 许可,重新发布在 huggingface.co/LibreYOLO
解读
Apache-2.0 is a permissive license, so this checkpoint can be used in commercial and closed-source products. It asks you to keep the license text and attribution notices with any copy of the weights you redistribute, and it grants a patent license. LibreYOLO vendors no Grounding DINO model source of its own: LibreGroundingDINO calls the Apache-2.0 `transformers` implementation, `GroundingDinoForObjectDetection`, directly, and downloads the official checkpoint into a LibreYOLO-hosted mirror repository that preserves the upstream snapshot files.

引用

@article{liu2023grounding,
  title={Grounding dino: Marrying dino with grounded pre-training for open-set object detection},
  author={Liu, Shilong and Zeng, Zhaoyang and Ren, Tianhe and Li, Feng and Zhang, Hao and Yang, Jie and Li, Chunyuan and Yang, Jianwei and Su, Hang and Zhu, Jun and others},
  journal={arXiv preprint arXiv:2303.05499},
  year={2023}
}

复制自作者在 github.com/IDEA-Research/GroundingDINO#black_nib-citation 上提供的引用块。

已针对 LibreYOLO v1.5.0 验证。本页的支持表、检查点和基准测试数据由已发布的库和权重生成,并非手工编写。