| Jun 09, 2026 | Co-authored “Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR” with M. Shi, K. Zhang, Z. Wang, N. B. Shankar, and A. Alwan, accepted to INTERSPEECH 2026. [arXiv] |
| Jun 01, 2026 | Co-authored “Compositional Domain Adaptation for Automatic Speech Recognition with Headwise Selective Attention Merging” with N. B. Shankar, Z. Wang, and A. Alwan, available online in Computer Speech & Language (vol. 101, 102012). [DOI] |
| Mar 01, 2026 | Our paper “Improving Zero-Shot Style Transfer Text-to-Speech by Disentangled Fine-Grained Style Modeling” appeared in JASA Express Letters 6(3), 034802. It is a lightweight (48M-parameter) style-transfer TTS system reaching a 3.64 prosody pMOS score. |
| Feb 05, 2026 | STACodec, our work with K. Zhang, M. Shi, N. B. Shankar, Z. Wang, and A. Alwan on balancing acoustic fidelity and semantic information in neural audio codecs, was accepted to ICASSP 2026. [arXiv] |
| Aug 12, 2025 | Posted ProMode, a zero-shot speech prosody model that maps text and reference audio to latent pitch, duration, and energy via a Perceiver-IO architecture, accepted to INTERSPEECH 2025. [arXiv] |
| Jan 14, 2025 | Co-authored “Selective Attention Merging for Low Resource Tasks: A Case Study of Child ASR” with N. B. Shankar, Z. Wang, and A. Alwan, accepted to ICASSP 2025. [arXiv] |
| Apr 01, 2024 | Honored to be named a 2024 Amazon Fellow. My Ph.D. research is supported by the Amazon Graduate Fellowship in ECE at UCLA. |