LibreVLA

LibreVLA runs robot policies from camera frames and robot state and returns action chunks.

Tasks
robot policies
Sizes
smolvla: base at 512 px; act_policy: base at 224 px; diffusion_policy: base at 224 px
Install
pip install "libreyolo[vla]"
Support tier
Sibling tier, since v1.6.0. A separate product surface with its own factory and contract.
Licenses
Code MIT, weights Apache-2.0. Commercial use

Install

bash
pip install "libreyolo[vla]"

Predict

Python
from libreyolo import LibreVLA, SAMPLE_IMAGE # Python 3.12 or later. This demonstrates shapes with synthetic state.model = LibreVLA("smolvla-base", device="cpu")result = model.predict(SAMPLE_IMAGE, state=[0.0] * 6, instruction="pick up the object")print(result.actions.data.shape)print(result.actions.first)

The extra requires Python 3.12 or later. Supply real camera observations and the state representation used during training for meaningful actions. The library returns action values; the caller owns the robot control loop. Call reset() between episodes.

Train

SmolVLA

LibreVLA() selects SmolVLA. Training updates its action expert with the vision backbone frozen. Pass a LeRobot dataset ID or local v3 directory to train(data=...). Defaults are 5 epochs, batch 8, workers 0 and val_split=0.1.

ACT

LibreVLA("act") constructs an untrained action policy. Train it on a LeRobot dataset before prediction. It uses the dataset camera, state and action shapes without a language instruction. Its image backbone may download separately.

Diffusion Policy

LibreVLA("diffusion") also starts without pretrained policy weights. It keeps observation history between calls; reset() clears that history between episodes.

Validate

val(data=...) reports val/action_l1, val/action_mse, val/action_l1_first and per-dimension val/action_l1_dims. These compare predicted and recorded actions offline; they are not robot task success rates.

Checkpoints

Training saves directories containing policy weights, processors and libreyolo_vla.json. Reload a returned best-checkpoint directory with LibreVLA(path). See the policy API.

Licensing

Check the license on the Hugging Face repository of the specific weights you download. Every checkpoint in the LibreYOLO org carries one, and they are not always the same across a family. That repository is the authoritative source; the summary below describes what applied when this page was last verified.

This is a description of the licenses involved, not legal advice. If the answer matters commercially, read the licenses yourself and take your own counsel.

Original work
SmolVLA, ACT and Diffusion Policy, Hugging Face and policy authors
Upstream license
Apache-2.0
LibreYOLO code
MIT
Weights
Apache-2.0, distributed by their authors. LibreYOLO does not host or mirror them.
Interpretation
The upstream policy runtime and SmolVLA base declare Apache-2.0. ACT and Diffusion Policy training do not load pretrained policy weights.

Verified against LibreYOLO v1.6.0. Support tables, checkpoints and benchmark numbers on this page are generated from the released library and the published weights, not written by hand.