7+ years of experience in infrastructure engineering, systems architecture, or a senior technical role focused on large-scale infrastructure
Proven experience designing multi-cloud architectures spanning AWS and at least one other major cloud provider or on-premises environment
Deep expertise in storage system design -- block, object, and file storage, including performance tuning for large-scale data workloads
Strong experience with compute orchestration using Kubernetes, and an understanding of how to schedule diverse workloads efficiently
Hands-on experience with GPU infrastructure -- procurement considerations, cluster design, driver and runtime management
Track record of capacity planning and infrastructure scaling for high-growth environments
Ability to communicate complex architectural decisions clearly to both technical and non-technical stakeholders
Strong understanding of networking fundamentals as they relate to infrastructure architecture (see our Network Engineer role for the deep specialist)
Direct experience architecting infrastructure for ML training workloads -- distributed training, large dataset management, experiment infrastructure
Background in cost optimization and FinOps practices for large-scale cloud and bare metal infrastructure
Experience operating and managing bare metal infrastructure in colocation facilities
Expertise in network architecture design, including high-bandwidth GPU interconnects and global traffic routing
Experience with infrastructure modeling and simulation for capacity planning
Familiarity with Slurm, Ray, or other HPC/ML job scheduling systems
Understanding of power, cooling, and physical infrastructure considerations for GPU-dense deployments
Notice: We're aware of individuals impersonating Deepgram recruiters. All legitimate Deepgram recruiting communication comes from an .com email address. If you've received a message claiming to be Deepgram, please forward it to careers .com