Have 6+ years of experience in software engineering, with a track record of owning significant technical scope within a team (e.g., driving a project from design through production, or acting as a de facto tech lead on a workstream)
Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
Solid grasp of distributed systems fundamentals — fault tolerance, graceful degradation, and failure handling in large-scale environments
Experience operating the control plane and low-level pieces of large-scale Kubernetes clusters
Experience with observability at scale: Prometheus, Grafana, distributed tracing, and building actionable alerting systems
Strong programming skills in Go and Python; ability to collaborate effectively on shared codebases
Solid knowledge of Linux systems, networking, containers, and cloud infrastructure
Take pride in owning and delivering core components of products and platforms
Experience building and operating managed Kubernetes services (GKE, EKS, AKS, or similar) or working on Kubernetes control plane components
Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning, or similar
Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes
Exposure to storage architecture for AI/ML workloads
Past contributions to CNCF projects or Kubernetes SIGs a plus
If you don’t meet all of these requirements but believe you may be a good fit, please still apply and provide a cover letter that helps us understand your experience and readiness for this role
Why Lambda
Lambda is building the essential infrastructure for the AI era. We're not just another cloud provider: we're a company founded by ML practitioners, for ML practitioners. Our customers include leading AI research labs and enterprises pushing the boundaries of what's possible with artificial intelligence