You are a creative, innovative engineer who operates at high velocity. You don't just solve problems. You find elegant solutions and ship them quickly. You embrace modern tools and AI-assisted development (like Claude Code) to accelerate your productivity and multiply your impact. You're energized by building new things, not maintaining the status quo
10+ years of experience in software engineering, platform engineering, or SRE, with at least 5 years focused on Kubernetes at scale
Expert-level understanding of Kubernetes internals: API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
Holistic infrastructure expertise: you've synthesized knowledge across compute, networking, storage, and security, not just Kubernetes in isolation. You can build solutions that span the full stack
Strong software engineering skills in Go (required) and Python; you write production-quality code, not just scripts
Deep experience with GPU orchestration in Kubernetes: NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling. Familiarity with NVIDIA Network Operator and GPUDirect is strongly preferred
Proven track record of technical leadership: driving design decisions across teams, mentoring engineers, and influencing infrastructure direction beyond your immediate scope
Deep experience designing and operating managed services or multi-tenant platforms. You understand what it takes to run infrastructure for external customers
Strong understanding of distributed systems principles: consensus, fault tolerance, consistency models, and graceful degradation
Experience with observability at scale: Prometheus, Grafana, distributed tracing, and building actionable alerting systems
Solid knowledge of Linux systems and networking (L2-L7), including high-performance networking concepts (RDMA, InfiniBand, RoCE)
Experience with infrastructure-as-code and GitOps workflows
Experience building and operating managed Kubernetes services (GKE, EKS, AKS, or similar) or working on Kubernetes control plane components
Hands-on experience with NVIDIA's open-source ecosystem beyond GPU Operator: Network Operator, NCCL tuning, Topograph, AICR, or similar emerging projects
Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
Background in confidential computing
Experience migrating customers or workloads from legacy/bespoke infrastructure to standardized platforms
Contributions to CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects
Familiarity with security and compliance in multi-tenant environments: RBAC, Pod Security Standards, network policies, workload isolation
Background in ML infrastructure: training clusters, inference serving, simulation
Why Lambda
Lambda is building the essential infrastructure for the AI era. We're not just another cloud provider: we're a company founded by ML practitioners, for ML practitioners. Our customers include leading AI research labs and enterprises pushing the boundaries of what's possible with artificial intelligence