Z-sight
Unifying Vision-Language-Action and Latent World Modeling
Anonymous ICLR 2027 submission · supplementary videos
Z-sight is a unified model that predicts the future in the native visual token space of its own vision-language model, the space in which it perceives the present, and generates actions while reading that prediction. Videos are played at real time unless noted otherwise.
Real-world tasks
Z-sight on the five evaluation tasks on a Unitree G1 humanoid.
Comparison with baselines
The same task and initial configuration for each policy.
The policy acts on its predicted future
The same trained model on Dog, with the foresight stream intervened at test time. When the prediction says the scene will not change, the robot still reaches toward the toy but does not grasp it or lift it into the box.