Experience with AI infrastructure, GPU platforms, or high-performance computing
Experience operating distributed systems across multiple regions or data centers
Experience with Kubernetes controllers, operators, CRDs, admission control, or scheduler extensions
Experience with etcd performance, backup, restore, or disaster recovery
Experience with Linux systems, container runtimes, cgroups, storage, or networking
Experience with chaos engineering, fault injection, or automated remediation
Experience with Kubernetes RBAC, OIDC, workload identity, or certificate management
Familiarity with SOC 2, ISO 27001, or similar compliance frameworks
Salary Range Information
Have 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering
Have deep experience operating Kubernetes in production
Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes
Have experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services
Are proficient with Terraform or similar infrastructure-as-code tools
Have built CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize
Have experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog
Can build production-quality tooling in Go, Python, or a similar language
Understand distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure
Have experience defining and operating against SLIs and SLOs
Can lead effectively during high-severity incidents
Approach recurring operational issues as engineering and automation problems
Communicate clearly and work effectively across teams
Bring strong ownership, sound judgment, and low ego