Senior-level experience in SRE, DevOps, or Platform Engineering. You've built and operated production infrastructure at scale and can own a significant workstream end to end with limited oversight
An active, opinionated use of AI development tools (Claude Code, Codex, etc.) in your own infrastructure workflow: Terraform changes, Kubernetes debugging, automation, operational investigations. You have a point of view on where these tools help and where they don't
Deep Kubernetes (EKS) and AWS experience (IAM, VPC, ECR, SSM/Secrets Manager, S3, SQS, Lambda, RDS/Aurora)
Strong IaC (Terraform) and GitOps experience, including PR-driven apply workflows (Atlantis or similar) and ArgoCD
CI/CD depth (GitHub Actions), including caching/parallelism, artifact management, test reliability, and pipeline observability
Full incident-ownership experience: you've carried on-call for systems you built and driven incidents from detection through follow-up
Excellent cross-team communication: you can translate platform constraints into developer-friendly solutions and documentation
Identity, access, and policy-as-code experience: workload/service identity (SPIFFE/SPIRE, OIDC), short-lived credentials, secrets management (Vault), and policy enforcement
Experience building self-service developer platforms, ephemeral environments, CLIs, scaffolding tools, or internal developer portals (Backstage or custom). You treat infrastructure as a product
Cost-awareness: you've built cost attribution, budgets, or rightsizing into a platform
Our Tech Stack
Infrastructure: AWS, EKS, Terraform (with Atlantis), Vault, Docker, OPA (Open Policy Agent)
CI/CD: GitHub Actions, ArgoCD + Kustomize (GitOps)
Messaging: Kafka (Confluent Cloud)
Observability: Datadog, OpenTelemetry
Languages/Apps: Node.js/TypeScript microservices, Python, React front-ends