BL
Brian Li
Member of Technical Staff (Singapore)
SingaporeAMI Labs
← Back to Team
Biography

Brian Li is a machine learning researcher and engineer specializing in multimodal artificial intelligence, currently serving as Member of Technical Staff (Singapore).

Prior to joining AMI Labs, Li held a position as Staff Research Scientist at ByteDance, where he was a core member of the Seed Multimodal and World Model team. During his tenure at ByteDance, Li made significant contributions to the field of large multimodal models as a co-author of several influential research works, including LLaVA-NeXT, LLaVA-OneVision, and LLaVA-Video.

These works are part of the LLaVA (Large Language and Vision Assistant) family of open-source multimodal models, which have been widely adopted by the research community and have contributed substantially to advances in vision-language understanding.

Before his role at ByteDance, Li also gained research experience at Microsoft Research, further grounding his expertise in large-scale AI systems and applied research.

Li's research focus centers on multimodal learning, vision-language models, and video understanding — areas that sit at the intersection of computer vision and natural language processing. His co-authorship on the LLaVA series of papers represents some of the most cited and reproduced work in open multimodal AI research in recent years, reflecting his ability to contribute to both the scientific and engineering dimensions of large-scale model development.

Career History
2022-2024
ByteDance (Seed Multimodal)
Staff Research Scientist
2020-2022
Microsoft Research
Research Scientist
Key Papers
We present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series. Our experimental results demonstrate that LLaVA-OneVision is the first single model that can simultaneously push the performance boundaries of open LMMs in three important computer vision scenarios: single-image, multi-image, and video scenarios. Importantly, the design of LLaVA-OneVision allows strong transfer learning across different modalities/scenarios, yielding new emerging capabilities. In particular, strong video understanding and cross-scenario capabilities are demonstrated through task transfer from images to videos.
2024 · arXiv
2,893 citations
More Publications (top 5 by citations)
2024 · Trans. Mach. Learn. Res.
2,986 citations
2024 · Trans. Mach. Learn. Res.
503 citations
2024 · North American Chapter of the Association for Computational Linguistics
346 citations