DreamerV3
Danijar Hafner and collaborators
A model-based reinforcement-learning reference: learn dynamics, then train behavior in imagined trajectories.
What you can explore
Experience containing observations, actions, rewards and episode boundaries. Predicted latent states and rewards; an actor-critic policy supplies actions.
Before you use it
Diverse control benchmarks, not a demonstrated driver for every dexterous hand.
Original contributors
Danijar Hafner and collaborators
License / access: MIT repository code; environment assets may have other terms.
Go deeper into this methodInputs, outputs, evidence, compute and release limits →
Editorial reference · checked 2026-09-25
