Z-sight

Unifying Vision-Language-Action and Latent World Modeling

Anonymous ICLR 2027 submission · supplementary videos

Z-sight is a unified model that predicts the future in the native visual token space of its own vision-language model, the space in which it perceives the present, and generates actions while reading that prediction. Videos are played at real time.

Real-world tasks

Z-sight on real-world tasks on a Unitree G1 humanoid.

DogPlace the toy dog into the box.
PuffsPick up the Puffs box, turn around and put it in the drawer.
ChairPush the chair under the table and return to standing.
BookCarry the book from the table to the shelf.

The policy acts on its predicted future

All videos are on Dog. The post-training variants are trained from the same checkpoint without future supervision or without the foresight branch. The test-time interventions keep the trained weights and replace the foresight stream. When the prediction says the scene will not change, the robot still reaches toward the toy but does not grasp it or lift it into the box.

Z-sight oursFull model.
No view lossPost-training variant: no future supervision (λv = 0).
No foresight branchPost-training variant: foresight expert removed.
Copy of current frameTest-time intervention: prediction says nothing will change.
NoiseTest-time intervention: no prediction available.