Build and fine-tune large scale ASR systems robust to accents, noise, telephony artifacts, and code switching
Leverage self-supervised pretraining and large-scale weak supervision
Improve transcription accuracy for real-world enterprise scenarios, including structured extraction and conversational nuance
Research and implement neural audio codecs that achieve extreme compression with minimal perceptual loss
Explore discrete and continuous latent representations for scalable speech modeling
Design codec architectures that enable downstream generative modeling and controllable synthesis
Design ablation studies that isolate the impact of architectural changes
Measure improvements using both objective metrics and perceptual evaluations
Validate ideas quickly through focused experiments that confirm or eliminate hypotheses
WHAT MAKES YOU A GREAT FIT
Track record of designing controlled experiments and meaningful ablations
Comfortable working with both offline benchmarks and live production metrics
Ability to move quickly from hypothesis to validation
Comfortable in fast-moving startup environments
Strong ownership mindset from research through deployment
Excited by ambiguous, unsolved problems
You treat unsolved problems as opportunities to invent new paradigms
You identify the single experiment that can validate an idea in days, not months
You measure everything and let data drive decisions
You are obsessed with making voice agents sound truly human
You use AI tools aggressively to amplify your own impact and accelerate research cycles