AI
Adithya Iyer
Member of Technical Staff
New YorkAMI Labs
← Back to Org Chart
Biography

Adithya Iyer is a Member of Technical Staff at AMI Labs, where he works on world modeling and advanced AI systems. His research interests span computer vision, generative AI, large language models and applied mathematics.

Before joining AMI Labs, Iyer was the founding machine-learning researcher at Morphic, where he led work on video-generation models. He previously conducted research on diffusion models and vision-language models at the NYU Center for Data Science under Professor Saining Xie, contributing to Cambrian-1, an open, vision-centric exploration of multimodal large language models presented at NeurIPS 2024. His earlier experience includes applied research at eBay and consulting at McKinsey & Company.

Iyer also co-founded Budnip, a startup that used deep learning and satellite imagery to automate geospatial analysis. The company was selected for the European Space Agency’s Copernicus Accelerator and placed third in the Copernicus Masters competition. He holds a master’s degree in Computational Science from New York University and studied at the Indian Institute of Technology Bombay, including a semester exchange at the Technical University of Denmark.

Career History
2024-2026
Morphic
Founding ML Researcher
2023-2024
NYU Center for Data Science
Research Assistant
Key Papers
Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging internet-scale monocular videos. Our core contribution is the generation of pseudo multi-view training triplets, consisting of a source video, a geometric anchor, and a target video. We achieve this by extracting distinct smooth random-walk crop trajectories from a single input video to serve as the source and target views. The anchor is synthetically generated by forward-warping the first frame of the source with a dense tracking field, which effectively simulates the distorted point-cloud inputs expected at inference. Because our independent cropping strategy introduces spatial misalignment and artificial occlusions, the model cannot simply copy information from the current source frame. Instead, it is forced to implicitly learn 4D spatiotemporal structures by actively routing and re-projecting missing high-fidelity textures across distinct times and viewpoints from the source video to reconstruct the target. At inference, our minimally adapted diffusion transformer utilizes a 4D point-cloud derived anchor to achieve state-of-the-art temporal consistency, robust camera control, and high-fidelity novel view synthesis on complex dynamic scenes.
2026 · Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
1 citations
We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision components are often insufficiently explored and disconnected from visual representation learning research. This gap hinders accurate sensory grounding in real-world scenarios. Our study uses LLMs and visual instruction tuning as an interface to evaluate various visual representations, offering new insights into different models and architectures—self-supervised, strongly supervised, or combinations thereof—based on experiments with over 15 vision models. We critically examine existing MLLM benchmarks, addressing the difficulties involved in consolidating and interpreting results from various tasks. To further improve visual grounding, we propose spatial vision aggregator (SVA), a dynamic and spatially-aware connector that integrates vision features with LLMs while reducing the number of tokens. Additionally, we discuss the curation of high-quality visual instruction-tuning data from publicly available sources, emphasizing the importance of distribution balancing. Collectively, Cambrian-1 not only achieves state-of-the-art performances but also serves as a comprehensive, open cookbook for instruction-tuned MLLMs. We provide model weights, code, supporting tools, datasets, and detailed instruction-tuning and evaluation recipes. We hope our release will inspire and accelerate advancements in multimodal systems and visual representation learning.
2024 · Advances in Neural Information Processing Systems
998 citations
More Publications (top 5 by citations)
2024 · Neural Information Processing Systems
905 citations
2023 · 2023 Intelligent Computing and Control for Engineering and Business Systems (ICCEBS)