Take one well-scoped problem from literature review through implementation, experimentation, and results
Design ablations that isolate what actually caused an improvement
Present your findings to the research team and defend the methodology
Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard
Use our distributed GPU infrastructure rather than toy-scale setups
Where the result warrants it, work with engineers to move it toward production
Choose your depth
Depending on your background and interests, your project may focus on
Expressive and controllable text-to-speech, including prosody and emotion modeling
Neural audio codecs and discrete or continuous speech representations
ASR robustness for telephony, accents, and code switching
Real-time and streaming inference under latency constraints
Full-duplex conversation and turn-taking dynamics
WHAT MAKES YOU A GREAT FIT
Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning
Strong intuition for audio quality and what makes synthetic speech sound wrong
Prior publications or open source contributions in speech or language AI are a strong signal, though not required
Fluent in PyTorch and comfortable in a real codebase
Able to run your own experiments on GPU clusters without waiting to be unblocked
You identify the single experiment that validates an idea in days, not months
You measure everything and let data drive decisions
You are honest about negative results, because they are how we narrow the search
You are obsessed with making voice agents sound truly human
You use AI tools aggressively to amplify your own impact