← Back to Org ChartBiography
Ayaan Naveed Malik is a research scientist at AMI Labs in New York, where he works on world modeling with Yann LeCun and Saining Xie. Before joining AMI, he was a research intern at NVIDIA’s GEAR Lab, where he developed the robot world-model projects DreamDojo and DreamZero under the guidance of Jim Fan, Yuke Zhu and Joel Jang. He previously conducted research at Together AI on enabling large language models to interpret visual information.
Malik is completing a bachelor’s degree in computer science, mathematics and philosophy at Stanford University. He previously founded Seekho, an education initiative focused on improving learning opportunities in South Asia, and contributed to the development of Pakistan’s standardized national curriculum. His academic distinctions include representing Pakistan on its national mathematics team and placing first in the NASA Ames Space Settlement Contest.
Career History
2022-2026
Stanford University
Bachelor's Degree Mathematics
Key Papers
State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions, using video as a dense representation of how the world evolves. By jointly modeling video and action, DreamZero learns diverse skills effectively from heterogeneous robot data without relying on repetitive demonstrations. This results in over 2x improvement in generalization to new tasks and environments compared to state-of-the-art VLAs in real robot experiments. Crucially, through model and system optimizations, we enable a 14B autoregressive video diffusion model to perform real-time closed-loop control at 7Hz. Finally, we demonstrate two forms of cross-embodiment transfer: video-only demonstrations from other robots or humans yield a relative improvement of over 42% on unseen task performance with just 10-20 minutes of data. More surprisingly, DreamZero enables few-shot embodiment adaptation, transferring to a new embodiment with only 30 minutes of play data while retaining zero-shot generalization.
2026 · arXivLabs
169 citations
More Publications (top 5 by citations)
2026 · arXiv.org
200 citations
2026 · arXiv.org
78 citations
2025 · BMJ Sexual & Reproductive Health
2025 · Journal of Health Inequalities
2025 · Folia Morphologica