BS
Bowen Shi
Member of Technical Staff
New York CityAMI Labs
← Back to Org Chart
Biography

Bowen Shi is a Member of Technical Staff at AMI Labs in New York, where he works on world models. His research sits at the intersection of speech, audio, vision and multimodal machine learning, with a focus on creating systems that can understand and generate information across different forms of media.

Before joining AMI Labs, Shi spent nearly four years as a Staff Research Scientist at Meta Superintelligence Labs and Facebook AI Research. He was a core contributor to Meta’s speech and audio foundation-model efforts, including SAM Audio, which extends the Segment Anything approach to general-purpose audio separation; MovieGen Audio; AudioBox; VoiceBox; and MMS, a project that scaled speech technology to more than 1,000 languages. His work also included multimodal representation learning and large-scale audio-generation and assessment systems.

Shi earned a PhD in computer science from the Toyota Technological Institute at Chicago, where he researched automatic sign-language understanding under Professor Karen Livescu. He also holds a master’s degree in computer science from Université Pierre et Marie Curie, an engineering degree from ENSTA Paris and a bachelor’s degree in mechatronics from Shanghai Jiao Tong University.

Career History
2023-2026
Meta
Staff Research Scientist
2021
Facebook AI
Research Intern
Key Papers
General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed categories such as speech or music, or limited in controllability, supporting only a single prompting modality such as text. In this work, we present SAM Audio, a foundation model for general audio separation that unifies text, visual, and temporal span prompting within a single framework. Built on a diffusion transformer architecture, SAM Audio is trained with flow matching on large-scale audio data spanning speech, music, and general sounds, and can flexibly separate target sources described by language, visual masks, or temporal spans. The model achieves state-of-the-art performance across a diverse suite of benchmarks, including general sound, speech, music, and musical instrument separation in both in-the-wild and professionally produced audios, substantially outperforming prior general-purpose and specialized systems. Furthermore, we introduce a new real-world separation benchmark with human-labeled multimodal prompts and a reference-free evaluation model that correlates strongly with human judgment.
2025 · arXivLabs