demos
Figures and audio samples from projects I led. Speech is best judged by ear, so most of these let you compare my systems against the baselines directly.
ProMode
A zero-shot prosody model that predicts pitch, duration and energy from text plus a short reference clip, then makes downstream TTS sound more natural.
Fine-Grained Style Transfer TTS
A lightweight (48M-parameter) zero-shot style-transfer TTS system that disentangles speaker from style, reaching 3.64 prosody pMOS against much larger baselines.
FusedF0
A hybrid DSP + DNN pitch tracker that stays accurate in noise by fusing summary-correlograms with raw waveform representations.
Speaker-Adaptive Postfiltering
Few-shot neural postfilters that adapt embedded text-to-speech to a new speaker from very little data. This work is the core of my M.S. thesis.
Uncertainty for Anti-Spoofing
Bayesian neural networks that say how confident they are when deciding whether a speaker-verification attempt is spoofed.