Efficient on-device audio
Wake-word and small-footprint audio architectures designed for strict compute and memory budgets, including query-by-example keyword spotting and BC-ResNet.
About
I am Byeonggeun Kim (김병근), a Senior Applied Scientist at Amazon AGI. My work follows three connected lines: efficient on-device audio, few-shot and robust learning, and audio foundation and generative models.
Background
Across these shifts, the goal has remained consistent: build audio systems that work under real constraints, from compute and limited data to latency and scale.
At Amazon AGI, my earlier work spanned audio encoders, neural codecs, audio and music generation, and multimodal LLM integration, including contributions to Amazon Nova. I now focus on full-duplex speech-to-speech modeling. Before Amazon, I developed efficient on-device and data-efficient learning methods at Qualcomm AI Research. I earned my M.S. at KAIST, advised by Soo-Young Lee.
Research trajectory
Wake-word and small-footprint audio architectures designed for strict compute and memory budgets, including query-by-example keyword spotting and BC-ResNet.
Few-shot, open-set, and domain-robust learning for limited labels and device shift. This line overlaps on-device research through small-footprint KWS and DCASE 2021.
Audio encoders, neural codecs, generative audio, and multimodal systems across Amazon AGI projects, including Amazon Nova. Current work centers on real-time, full-duplex speech-to-speech interaction.
Experience
Seattle, WA
Audio and Speech AGI
Toronto, ON & Seattle, WA
Earlier work spanned audio and music generation, audio understanding encoders, and neural codecs, including contributions to Amazon Nova. Current work focuses on full-duplex speech-to-speech systems.
Seoul, South Korea
Developed efficient on-device and data-efficient learning methods for wake-word and acoustic scene recognition, including BC-ResNet, few-shot keyword spotting, and the DCASE 2021 first-place system.
Recognition & talks
Efficient and High Fidelity Text-to-Audio Generation: From Parallel Decoding to Causal Streaming.
Invited talk on autoregressive text-to-audio generation approaches.
Q3/Q4 award for contributions to Large Language Models raising the functional bar.
View award photoDomain adaptation for device-imbalanced acoustic scene classification.
Education
KAIST · Advisor: Soo-Young Lee
KAIST · Double major in Business