3D-MOOD

3D-MOOD predicts three-dimensional boxes from a single image.

Tasks
3d detection
Install
pip install "libreyolo[hf]"
Support tier
Sibling tier, since v1.6.0. A separate product surface with its own factory and contract.
Upstream
3D-MOOD by ETH Zurich, Computer Vision and Geometry Lab, Apache-2.0. Paper, source
Licenses
Code MIT, weights Apache-2.0. Commercial use

Install

bash
pip install "libreyolo[hf]"

Predict

Python
from libreyolo import Libre3DMOOD, SAMPLE_IMAGEimport numpy as npfrom PIL import Image # Requires the separately installed upstream runtime.model = Libre3DMOOD(device="cpu")# Rough 3x3 pinhole guess; replace with measured calibration for the original image.w, h = Image.open(SAMPLE_IMAGE).sizeintrinsics = np.array([[w, 0, w / 2], [0, w, h / 2], [0, 0, 1]], dtype=np.float32)result = model.predict(SAMPLE_IMAGE, intrinsics=intrinsics, text="person")print(result.boxes3d)

3D-MOOD performs open-set detection using text labels and original-image camera intrinsics. Install its separate upstream runtime and provide runtime_path and runtime_python when needed.

result.boxes3d holds centers, dimensions and wxyz quaternions in camera coordinates, plus confidence, class and intrinsics. Plotting projects cuboids into the image. Training, validation, tracking and export are not supported. See 3D detection.

Licensing

Check the license on the Hugging Face repository of the specific weights you download. Every checkpoint in the LibreYOLO org carries one, and they are not always the same across a family. That repository is the authoritative source; the summary below describes what applied when this page was last verified.

This is a description of the licenses involved, not legal advice. If the answer matters commercially, read the licenses yourself and take your own counsel.

Original work
3D-MOOD, ETH Zurich, Computer Vision and Geometry Lab
Upstream license
Apache-2.0
Upstream source
github.com/cvg/3D-MOOD
LibreYOLO code
MIT
Weights
Apache-2.0, republished at huggingface.co/LibreYOLO
Interpretation
The source and model publisher declare Apache-2.0.

Citation

@InProceedings{Yang_2025_ICCV,
    author    = {Yang, Yung-Hsu and Piccinelli, Luigi and Segu, Mattia and Li, Siyuan and Huang, Rui and Fu, Yuqian and Pollefeys, Marc and Blum, Hermann and Bauer, Zuria},
    title     = {3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection},
    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
    month     = {October},
    year      = {2025},
    pages     = {7429-7439}
}

Copied from the authors' citation block at raw.githubusercontent.com/cvg/3D-MOOD/main/README.md.

Verified against LibreYOLO v1.6.0. Support tables, checkpoints and benchmark numbers on this page are generated from the released library and the published weights, not written by hand.