Z-sight
Unifying Vision-Language-Action and Latent World Modeling
Z-sight is a unified model that predicts the future in the native visual token space of its own vision-language model, the space in which it perceives the present, and generates actions while reading that prediction. Videos are played at real time.
Real-world tasks
Z-sight on real-world tasks on a Unitree G1 humanoid.
The policy acts on its predicted future
All videos are on Dog. The post-training variants are trained from the same checkpoint without future supervision or without the foresight branch. The test-time interventions keep the trained weights and replace the foresight stream. When the prediction says the scene will not change, the robot still reaches toward the toy but does not grasp it or lift it into the box.