Bowen Shi is a Member of Technical Staff at AMI Labs in New York, where he works on world models. His research sits at the intersection of speech, audio, vision and multimodal machine learning, with a focus on creating systems that can understand and generate information across different forms of media.
Before joining AMI Labs, Shi spent nearly four years as a Staff Research Scientist at Meta Superintelligence Labs and Facebook AI Research. He was a core contributor to Meta’s speech and audio foundation-model efforts, including SAM Audio, which extends the Segment Anything approach to general-purpose audio separation; MovieGen Audio; AudioBox; VoiceBox; and MMS, a project that scaled speech technology to more than 1,000 languages. His work also included multimodal representation learning and large-scale audio-generation and assessment systems.
Shi earned a PhD in computer science from the Toyota Technological Institute at Chicago, where he researched automatic sign-language understanding under Professor Karen Livescu. He also holds a master’s degree in computer science from Université Pierre et Marie Curie, an engineering degree from ENSTA Paris and a bachelor’s degree in mechatronics from Shanghai Jiao Tong University.