Brian Li is a machine learning researcher and engineer specializing in multimodal artificial intelligence, currently serving as Member of Technical Staff (Singapore).
Prior to joining AMI Labs, Li held a position as Staff Research Scientist at ByteDance, where he was a core member of the Seed Multimodal and World Model team. During his tenure at ByteDance, Li made significant contributions to the field of large multimodal models as a co-author of several influential research works, including LLaVA-NeXT, LLaVA-OneVision, and LLaVA-Video.
These works are part of the LLaVA (Large Language and Vision Assistant) family of open-source multimodal models, which have been widely adopted by the research community and have contributed substantially to advances in vision-language understanding.
Before his role at ByteDance, Li also gained research experience at Microsoft Research, further grounding his expertise in large-scale AI systems and applied research.
Li's research focus centers on multimodal learning, vision-language models, and video understanding — areas that sit at the intersection of computer vision and natural language processing. His co-authorship on the LLaVA series of papers represents some of the most cited and reproduced work in open multimodal AI research in recent years, reflecting his ability to contribute to both the scientific and engineering dimensions of large-scale model development.