5+ years in cloud infrastructure, DevOps, SRE, or platform operations
Hands-on AWS experience: VPCs, EC2, S3, IAM, CloudWatch, Route 53, load balancers, security groups, private networking
Proficiency with IaC tooling (Terraform strongly preferred)
Strong Linux fundamentals — networking, process management, storage, troubleshooting
Experience with CI/CD, Git-based workflows, and monitoring/alerting platforms
Clear communicator who can document infrastructure and collaborate across engineering teams
Experience with AI/ML, GPU, or HPC workloads
Kubernetes on AWS (EKS or self-managed)
Observability platforms: Prometheus, Grafana, Loki, OpenTelemetry, Datadog
AWS cost optimization: right-sizing, savings plans, lifecycle policies, tagging
Startup or high-growth infrastructure environment background