Strong software engineering fundamentals, with proficiency in Python and experience writing production-quality, well-tested ML code
Hands-on experience taking ML models from research or prototype stage into production at scale — not just training models, but shipping and operating them
A working understanding of the modern deep learning stack (e.g., PyTorch) and the realities of training, evaluating, and serving large models
Experience building ML pipelines and tooling — training orchestration, evaluation harnesses, model packaging, deployment, or CI/CD for models
Familiarity with serving and inference optimization — latency, throughput, batching, and resource efficiency for production model workloads
Comfort operating across distributed systems and GPU compute, whether in the cloud, on bare metal, or both
A collaborative, builder mindset — you can partner with researchers, scope an ambiguous problem, and drive it to a measurable result
Experience with the research-to-production handoff specifically — building the systems and conventions that let research and engineering iterate together quickly
Background in speech, audio, or other real-time/streaming ML domains
Experience designing automated model evaluation and release-gating systems, including regression detection across model versions
Familiarity with hybrid infrastructure spanning on-premise GPU clusters and cloud, and with workload orchestration across them
Experience with inference optimization techniques (quantization, distillation, compilation, or runtime tuning) for production serving
A track record of building internal platforms or developer-facing tooling that measurably improved how a team ships models
Notice: We're aware of individuals impersonating Deepgram recruiters. All legitimate Deepgram recruiting communication comes from an .com email address. If you've received a message claiming to be Deepgram, please forward it to careers .com