Gemma 4

Gemma 4 produces object detections from image and text input.

Tasks
detection
Sizes
e2b, e4b at 896 px
Install
pip install "libreyolo[vlm]"
Support tier
Sibling tier, since v1.6.0. A separate product surface with its own factory and contract.
Licenses
Code MIT, weights Apache-2.0. Commercial use

Install

bash
pip install "libreyolo[vlm]" "transformers>=5.10.0"

Predict

Python
from libreyolo import LibreVLM, SAMPLE_IMAGE model = LibreVLM("gemma-4-e2b", device="cpu")model.set_classes(["person", "building"])result = model(SAMPLE_IMAGE)print(result.boxes)

The E2B and E4B adapters parse boxes on a 0–1000 coordinate grid. The bare gemma-4 alias selects E4B. Install Transformers 5.10 or later. Training is not supported.

Licensing

Check the license on the Hugging Face repository of the specific weights you download. Every checkpoint in the LibreYOLO org carries one, and they are not always the same across a family. That repository is the authoritative source; the summary below describes what applied when this page was last verified.

This is a description of the licenses involved, not legal advice. If the answer matters commercially, read the licenses yourself and take your own counsel.

Original work
Gemma 4, Google
Upstream license
Apache-2.0
Upstream source
huggingface.co/google
LibreYOLO code
MIT
Weights
Apache-2.0, republished at huggingface.co/LibreYOLO
Interpretation
The Gemma 4 adapters reference Apache-2.0 snapshots; this declaration does not apply to other Gemma generations.

Verified against LibreYOLO v1.6.0. Support tables, checkpoints and benchmark numbers on this page are generated from the released library and the published weights, not written by hand.