AB
Andrew Brown
Member of Technical Staff
New YorkAMI Labs
← Back to Org Chart
Biography

Andrew Brown is a Member of Technical Staff at AMI Labs, where he works on advanced machine intelligence. His research spans generative models for vision, multimodal representation learning and retrieval, and the understanding of people and narratives in video. Before joining AMI, he spent more than three years as a computer-vision and machine-learning research scientist within Meta’s generative-AI organization.

Brown’s research has addressed targeted image editing, large-scale image retrieval and the analysis of people and stories across visual and audiovisual media. His publications include End-to-End Visual Editing with a Generatively Pre-Trained Artist, presented at ECCV 2022, and Smooth-AP, an ECCV 2020 method for improving large-scale image-retrieval training. He also co-authored work on movie-story retrieval, speaker recognition and automated face labelling in video archives, with the latter receiving the Best Student Paper award at MIPR 2021.

Brown earned both his DPhil in Engineering Science and his Master of Engineering from the University of Oxford. He completed his doctorate in computer vision and machine learning within Oxford’s Visual Geometry Group under Professor Andrew Zisserman. Earlier in his career, he developed machine-learning software for Bühler’s optical food-sorting systems, with his work subsequently incorporated into a new generation of the company’s machines.

Career History
2023-2026
Meta
Research Scientist (Computer Vision and Machine Learning)
2018-2023
University of Oxford
DPhil - Computer Vision and Machine Learning
2021-2022
PRO Unlimited @ Meta
Research Scientist Contractor
2021-2021
Meta
Research Scientist
Key Papers
We present Emu Video, a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We identify critical design decisions–adjusted noise schedules for diffusion, and multi-stage training–that enable us to directly generate high quality and high resolution videos, without requiring a deep cascade of models as in prior work. In human evaluations, our generated videos are strongly preferred in quality compared to all prior work–[Math Processing Error] vs. Google’s Imagen Video, [Math Processing Error] vs. Nvidia’s PYOCO, and [Math Processing Error] vs. Meta’s Make-A-Video. Our model outperforms commercial solutions such as RunwayML’s Gen2 and Pika Labs. Finally, our factorizing approach naturally lends itself to animating images based on a user’s text prompt, where our generations are preferred [Math Processing Error] over prior work.
2024 · European Conference on Computer Vision
340 citations
More Publications (top 3 by citations)
2023 · European Conference on Computer Vision
292 citations