Eray Eren
/eˈɾaj eˈɾen/ · ARPAbet: EH0 R AY1 EH0 R EH1 N
Ph.D. Candidate, Electrical & Computer Engineering, UCLA · SPAPL, advised by Prof. Abeer Alwan.
I am a Ph.D. candidate in Electrical & Computer Engineering at UCLA, working in the Speech Processing and Auditory Perception Laboratory (SPAPL) under the supervision of Prof. Abeer Alwan. I expect to graduate in December 2026.
My research is on expressive speech synthesis: how to model and transfer the prosody and speaking style of a voice from very little reference audio. Recent work includes ProMode, a zero-shot prosody model built on a Perceiver-IO architecture that maps text and reference audio to latent pitch, duration, and energy. I also proposed a lightweight (48M-parameter) zero-shot style-transfer TTS system with fine-grained style disentanglement, which outperforms competitive baselines with a 3.64 prosody pMOS score. Other ongoing work includes robust F0 estimation, neural audio codecs, and the use of synthetic TTS data for speech recognition.
My Ph.D. is supported by an Amazon Graduate Fellowship (2024 Amazon Fellow) and a Flawless AI PhD Research Grant (2023–2025). I also spent three consecutive summers as a Research Scientist Intern at Flawless AI, working on style-preserving neural speech editing.
Before UCLA, I completed an M.S. in Computer Science at Ozyegin University (GPA 4.0/4.0) with Prof. Cenk Demiroglu, working on speaker-adaptive neural postfilters for embedded TTS and Bayesian uncertainty estimation for anti-spoofing in speaker verification. I hold a B.S. in Electronics Engineering from Bogazici University, and previously worked as an R&D engineer on Turkish text-to-speech systems and embedded software for home appliances.
Feel free to reach out if you would like to talk about speech synthesis, prosody, or generative audio models.
news
| Jun 09, 2026 | Co-authored “Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR” with M. Shi, K. Zhang, Z. Wang, N. B. Shankar, and A. Alwan, accepted to INTERSPEECH 2026. [arXiv] |
|---|---|
| Jun 01, 2026 | Co-authored “Compositional Domain Adaptation for Automatic Speech Recognition with Headwise Selective Attention Merging” with N. B. Shankar, Z. Wang, and A. Alwan, available online in Computer Speech & Language (vol. 101, 102012). [DOI] |
| Mar 01, 2026 | Our paper “Improving Zero-Shot Style Transfer Text-to-Speech by Disentangled Fine-Grained Style Modeling” appeared in JASA Express Letters 6(3), 034802. It is a lightweight (48M-parameter) style-transfer TTS system reaching a 3.64 prosody pMOS score. |
| Feb 05, 2026 | STACodec, our work with K. Zhang, M. Shi, N. B. Shankar, Z. Wang, and A. Alwan on balancing acoustic fidelity and semantic information in neural audio codecs, was accepted to ICASSP 2026. [arXiv] |
| Aug 12, 2025 | Posted ProMode, a zero-shot speech prosody model that maps text and reference audio to latent pitch, duration, and energy via a Perceiver-IO architecture, accepted to INTERSPEECH 2025. [arXiv] |
selected publications
- JASA Express Letters, 2026
- In Proc. INTERSPEECH, 2025
- Computer Speech & Language, 2023