Masked Visual Actions: A Pixel-Space Control Interface for Unified World Modeling

· Editorial Team estimated
world model visual actions pixel-space control robot manipulation inverse modeling

This paper introduces Masked Visual Actions, a pixel-space control interface that expresses robot actions as partially revealed trajectories of an arbitrary entity in a video. Finetuned with only 15 hours of masked examples, a single checkpoint achieves strong visual fidelity, forward dynamics prediction, inverse modeling, and model-based planning across diverse scenes and multiple embodiments.

Paper · arXiv:2607.19343

Background

Video models absorb rich priors about how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is communicating action to these models in a form aligned with the visual space in which they learned interaction priors, yet still grounded in physical manipulation. Prior approaches require either extensive action annotations or are restricted to specific robot embodiments.

Core Innovation

The paper introduces Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model predicting the scene’s response; revealing desired object motion makes the same model recover robot behavior consistent with that outcome. This symmetry enables a single model to support forward prediction, inverse modeling, and model-based planning without architectural modifications.

Key Results

Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation. It improves decision-making by ranking candidate futures in model-based planning and supports inverse modeling by synthesizing robot motion from desired object motion.

Limitations

The current approach requires masking trajectory examples for each scene, and generalization to entirely unseen embodiments or manipulation types needs further validation. While 15 hours of training data is remarkably efficient, the fidelity boundaries on extremely complex scenes remain unexplored.

Industry Implications

This approach offers a highly scalable path for robotic world modeling — leveraging the rich visual priors already present in video models with minimal action supervision. The ability to reuse a single checkpoint across embodiments and tasks could dramatically reduce data requirements for robot learning, accelerating skill acquisition and policy deployment across diverse robotic platforms.