Z-sight

Unifying Vision-Language-Action and Latent World Modeling

Anonymous ICLR 2027 submission · supplementary videos

Z-sight is a unified model that predicts the future in the native visual token space of its own vision-language model, the space in which it perceives the present, and generates actions while reading that prediction. Videos are played at real time unless noted otherwise.

Real-world tasks

Z-sight on the five evaluation tasks on a Unitree G1 humanoid.

DogPlace the toy dog into the box.
PuffsPick up the Puffs box, turn around and put it in the drawer.
ChairPush the chair under the table and return to standing.
BookCarry the book from the table to the shelf.
FruitPick up the plate, open the fridge, place the plate inside and close the door.

Comparison with baselines

The same task and initial configuration for each policy.

Z-sight ours
π0.5
GR00T N1.7

The policy acts on its predicted future

The same trained model on Dog, with the foresight stream intervened at test time. When the prediction says the scene will not change, the robot still reaches toward the toy but does not grasp it or lift it into the box.

Predicted futureFull model.
NoiseNo prediction available.
Copy of current framePrediction: nothing will change.