OpenVLA
OpenVLA research team
An instruction-conditioned action model. A useful example for understanding how a vision-language backbone becomes a robot policy.
What you can explore
Camera image and language instruction. Tokenized actions decoded with embodiment-specific normalization; the original example uses a 7-dimensional action.
Before you use it
Author-reported arm-and-gripper experiments, including WidowX and Google Robot. This is not a verified multi-finger hand controller.
Original contributors
OpenVLA research team
License / access: MIT code; released Llama-2-derived weights also carry base-model terms. Dataset terms remain separate.
Code, models and licensing notes
Editorial reference · checked 2026-09-25
