One of the first interpretability analyses of how video encoders (V-JEPA 2) represent physical variables.

Authors: Sonia Joseph, Quentin Garrido, Randall Balestriero, Matthew Kowal, Thomas Fel, Shahab Bakhtiari, Blake Richards, and Mike Rabbat.

Work done at Meta Superintelligence Labs.

Full preprint: https://arxiv.org/abs/2602.07050

In the first impossible example, a cube unexpectedly appears halfway through the video. In the second, a red cube disappears behind a wall and does not reappear. In the third, a ball reverses trajectory mid-flight. Each example violates a symmetry or continuity constraint that we would normally expect in the physical world.

This research moves the debate over world models from what they can do to how they represent the physical world.

By examining V-JEPA 2 layer by layer, the researchers identify a “Physics Emergence Zone” roughly one-third of the way through the network, where the model begins integrating motion across time and distinguishing physically possible events from impossible ones.

Crucially, the most useful representations appear in the middle of the model rather than at its output. The results offer rare mechanistic evidence that a system trained to learn from video can develop structured representations of physical dynamics without being explicitly taught Newtonian rules.

For the AMI, the broader significance lies in what this suggests about the path beyond language-centric AI. If world models are to support robots and other systems that can understand, predict and act in the physical world, benchmark performance alone will not be enough: their internal representations must also be interpretable, auditable and controllable.

This study shows both the promise and the difficulty. The models develop strikingly brain-like population codes for motion direction, yet those variables are distributed across dozens of dimensions rather than stored as simple, editable concepts. Understanding this internal geometry could therefore become essential to building safer physical agents, more reliable scientific simulators, and AI systems grounded in the dynamics of the real world.

Read a more detailed explainer from researcher Sonia Joseph.

↓ Open PDF in new tab