Amazon Nova 2
Audio and music generation, encoder, and codec contributions to Amazon's multimodal reasoning and generation models.
Senior Applied Scientist Amazon AGI
I build models that understand and generate audio for natural, real-time conversations.
My research has progressed from efficient on-device and few-shot audio learning to audio encoders, neural codecs, and generation for Amazon AGI projects, including Amazon Nova. I currently focus on full-duplex speech-to-speech modeling.
Selected work
Representative outcomes across efficient architecture, robust learning, and audio foundation models.
Audio and music generation, encoder, and codec contributions to Amazon's multimodal reasoning and generation models.
Generative audio language modeling with masked next-token prediction and continuous-valued tokens.
Broadcasted residual learning for accurate, low-compute keyword spotting on resource-constrained devices.
A low-complexity, device-robust acoustic scene classification system combining efficient design and domain generalization.
Recent
Invited talk at Yale University (CPSC 7760): Efficient and High Fidelity Text-to-Audio Generation.
Contributed to the launch of Amazon Nova 2.
Invited talk at the Amazon Media & Entertainment ML+AI Summit.
Two papers accepted at ICML 2025, including IMPACT and continuous-valued audio language modeling.