QG
Quentin Garrido
Member of Technical Staff
ParisAMI Labs
← Back to Org Chart
Biography

Quentin Garrido is a Member of Technical Staff at AMI Labs, where his research focuses on advanced machine intelligence. His work spans self-supervised learning, computer vision and world models, including the challenge of learning representations and intuitive physical structure directly from raw images and video.

Before joining AMI Labs, Garrido was a research scientist at Meta FAIR, where he also completed a PhD in joint supervision with Université Gustave Eiffel under the supervision of Yann LeCun and Laurent Najman. His research has examined the relationship between contrastive and non-contrastive learning, methods for evaluating self-supervised representations without labels, equivariant representation learning, and the use of world models in visual learning. His publications include RankMe, presented orally at ICML 2023, and a study of contrastive and non-contrastive self-supervised learning that received an Outstanding Paper Honorable Mention at ICLR 2023.

Garrido holds a PhD in computer science from Université Gustave Eiffel, a master’s degree from the Mathematics, Vision and Learning program at ENS Paris-Saclay, and an engineering degree in computer science from ESIEE Paris, where he graduated first in his class. His earlier work applied deep learning and graph-based methods to biological data, earning him the 2022 Ian Lawson Van Toch Memorial Award for an outstanding student paper on visualizing cellular development hierarchies in single-cell RNA sequencing data.

Career History
2024-2026
Meta
Research Scientist
Key Papers
This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K.
2024 · arXivLabs
517 citations
More Publications (top 5 by citations)
2026 · arXiv.org
5 citations