AB
Amir Bar
Member of Technical Staff
LondonAMI Labs
โ† Back to Org Chart
Biography

Amir Bar is an assistant professor at Imperial College London, where he is establishing a new research lab, and a founding member of technical staff at AMI Labs. His work spans computer vision, self-supervised learning, embodied AI and world models, with the broader goal of creating machines that can understand their surroundings, anticipate what will happen next and use those predictions to plan and act.

Before joining Imperial and AMI Labs, Bar was a research scientist and postdoctoral researcher at Meta FAIR, where he worked with Yann LeCun. His recent research includes Navigation World Models, which uses generated video to simulate possible trajectories for planning, and work on visual representation learning, robotic manipulation and embodied control. Navigation World Models received a Best Paper Honorable Mention at CVPR 2025, while EB-JEPA received the Outstanding Paper Award at the ICLR 2026 World Models Workshop.

Bar completed his PhD through Tel Aviv University and the University of California, Berkeley, under the supervision of Amir Globerson and Trevor Darrell. Earlier in his career, he spent nearly six years at Zebra Medical Vision, progressing from machine-learning researcher to AI research lead and technical lead while developing computer-vision systems for detecting acute findings in medical scans.

Career History
2026-Present
Imperial College London
Assistant Professor
2025-2026
Meta
Research Scientist
2016-2022
Zebra Medical Vision Ltd
AI Tech Lead
Key Papers
We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First, a dense predictive loss uses a masking-based objective in which both visible and masked tokens contribute to the training signal, encouraging explicit spatial and temporal grounding. Second, deep self-supervision applies the self-supervised objective hierarchically across multiple intermediate encoder layers to improve representation quality. Third, multi-modal tokenizers enable unified training across images and videos. Finally, the model benefits from effective scaling in both model capacity and training data. Together, these design choices produce representations that are spatially structured, semantically coherent, and temporally consistent. Empirically, V-JEPA 2.1 achieves state-of-the-art performance on several challenging benchmarks, including 7.71 mAP on Ego4D for short-term object-interaction anticipation and 40.8 Recall@5 on EPIC-KITCHENS for high-level action anticipation, as well as a 20-point improvement in real-robot grasping success rate over V-JEPA-2 AC. The model also demonstrates strong performance in robotic navigation (5.687 ATE on TartanDrive), depth estimation (0.307 RMSE on NYUv2 with a linear probe), and global recognition (77.7 on Something-Something-V2). These results show that V-JEPA 2.1 significantly advances the state of the art in dense visual understanding and world modeling.
2026 ยท arXivLabs
48 citations