Research output

Publications and granted patents.

Work across efficient on-device audio, few-shot and robust learning, and audio foundation and generative models.

Featured

Full list

Publications.

* Equal contribution. Links open the publisher, paper, or project page.

2026
  • Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qingming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang INTERSPEECH 2026 (Oral) · mentored internship work
  • Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Kuan-Po Huang, Bo-Ru Lu, Byeonggeun Kim, Mihee Lee, Zalan Fabian, Renard Korzeniowski, Qingming Tang, Greg Ver Steeg, Hung-yi Lee, Chieh-Chi Kao, Chao Wang arXiv preprint, 2026 · mentored internship work
2025
  • Amazon Nova 2: Multimodal reasoning and generation models Amazon Artificial General Intelligence Amazon Technical Report, 2025
  • Unlocking Transfer Learning for Open-World Few-Shot Recognition Byeonggeun Kim*, Jun-Tae Lee*, Kyuhong Shim, Simyung Chang NeurIPS 2025 — Reliable ML Workshop
  • Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Shu-wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang, Huy Phan, Bo-Ru Lu, Harshavardhan Sundar, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao, Chao Wang ICML 2025 · mentored internship work
  • IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Kuan-Po Huang, Shu-wen Yang, Huy Phan, Bo-Ru Lu, Byeonggeun Kim, Sashank Macha, Qingming Tang, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao, Chao Wang ICML 2025 · mentored internship work
  • Effective Techniques for Scaling Audio Encoder Pretraining Byeonggeun Kim*, Andrew Bydlon*, Qingming Tang, Huy Phan, Chieh-Chi Kao, Tao Zhang, Chao Wang ICASSP 2025 (Oral)
2024
  • Cross-Triggering Issue in Audio Event Detection and Mitigation Huy Phan, Byeonggeun Kim, Vu Nguyen, Andrew Bydlon, Qingming Tang, Chieh-Chi Kao, Chao Wang ICASSP 2024
2023
  • Task-Agnostic Open-Set Prototype for Few-Shot Open-Set Recognition Byeonggeun Kim*, Jun-Tae Lee*, Kyuhong Shim, Simyung Chang ICIP 2023
  • Improving Small Footprint Few-shot Keyword Spotting with Supervision on Auxiliary Data Seunghan Yang, Byeonggeun Kim, Kyuhong Shim, Simyung Chang INTERSPEECH 2023 (Oral)
  • Scalable Weight Reparametrization for Efficient Transfer Learning Byeonggeun Kim*, Jun-Tae Lee*, Seunghan Yang, Simyung Chang ICASSP 2023 (Oral)
  • TTN: A Domain-Shift Aware Batch Normalization in Test-Time Adaptation Hyesu Lim, Byeonggeun Kim, Jaegul Choo, Sungha Choi ICLR 2023 · mentored internship work
2022
  • Dummy Prototypical Networks for Few-shot Open-set Keyword Spotting Byeonggeun Kim, Seunghan Yang, Inseop Chung, Simyung Chang INTERSPEECH 2022
  • Domain Generalization with Relaxed Instance Frequency-wise Normalization for Multi-device Acoustic Scene Classification Byeonggeun Kim, Seunghan Yang, Jangho Kim, Hyunsin Park, Juntae Lee, Simyung Chang INTERSPEECH 2022
  • Personalized Keyword Spotting through Multi-task Learning Seunghan Yang, Byeonggeun Kim, Inseop Chung, Simyung Chang INTERSPEECH 2022 (Oral)
2021
  • Broadcasted Residual Learning for Efficient Keyword Spotting (BCResNets) Byeonggeun Kim*, Simyung Chang*, Jinkyu Lee, Dooyong Sung INTERSPEECH 2021
  • Domain Generalization on Efficient Acoustic Scene Classification Using Residual Normalization Byeonggeun Kim, Seunghan Yang, Jangho Kim, Simyung Chang DCASE 2021 Workshop
  • QTI submission to DCASE 2021: Residual normalization for device imbalanced acoustic scene classification with efficient design Byeonggeun Kim, Seunghan Yang, Jangho Kim, Simyung Chang DCASE Challenge 2021 — 1st place winner
2019
  • Orthogonality Constrained Multi-Head Attention For Keyword Spotting Mingu Lee, Jinkyu Lee, Hye Jin Jang, Byeonggeun Kim, Wonil Chang, Kyuwoong Hwang IEEE ASRU 2019
  • Query-by-Example On-Device Keyword Spotting Byeonggeun Kim, Mingu Lee, Jinkyu Lee, Yeonseok Kim, Kyuwoong Hwang IEEE ASRU 2019

Intellectual property

Granted patents.

  • Acoustic event detection Quoc Huy Phan, Byeonggeun Kim, Andrew Thomas Bydlon, Qingming Tang, Chieh-Chi Kao, Chao Wang, Tien Vu Nguyen US 12,646,502 · 2 Jun 2026
  • Multi-task learning for personalized keyword spotting Seunghan Yang, Byeonggeun Kim, Inseop Chung, Simyung Chang US 12,347,439 · 1 Jul 2025
  • Relaxed instance frequency normalization for neural-network-based audio processing Byeonggeun Kim, Seunghan Yang, Hyunsin Park, Juntae Lee, Simyung Chang US 12,266,379 · 1 Apr 2025
  • Target Keyword Selection Wonil Chang, Jinseok Lee, Mingu Lee, Jinkyu Lee, Byeonggeun Kim, Dooyong Sung, Jaewon Choi, Kyu Woong Hwang US 12,039,968 · 16 Jul 2024
  • Task Agnostic Open-set Prototypes for Few-shot Open-set Recognition Byeonggeun Kim, Juntae Lee, Simyung Chang US 12,019,641 · 25 Jun 2024
  • Systems and methods of image processing based on gaze detection Hyunsin Park, Juntae Lee, Simyung Chang, Byeonggeun Kim, Jaewon Choi, Kyu Woong Hwang US 11,798,204 · 24 Oct 2023
  • On-device self training in two-stage wakeup system Young Mo Kang, Sungrak Yun, Kyu Woong Hwang, Hye Jin Jang, Byeonggeun Kim US 11,664,012 · 30 May 2023
  • Activating speech recognition based on hand patterns detected using plurality of filters Sungrack Yun, Young Mo Kang, Hye Jin Jang, Byeonggeun Kim, Kyu Woong Hwang US 11,437,031 · 6 Sep 2022
  • Method and apparatus for activating speech recognition Byeonggeun Kim, Young Mo Kang, Sungrack Yun, Kyu Woong Hwang, Hye Jin Jang US 11,205,433 · 21 Dec 2021