Proven experience managing high-traffic, large-scale production environments
Deep proficiency with major Cloud providers (AWS, Azure, or GCP)
Strong understanding of Kubernetes internals (CNI, Ingress, Network Policies, etc.)
Hands-on experience IaC and automation tools (Terraform, Ansible, Helm, etc)
Deep understanding of Linux environments, including system internals and performance troubleshooting
Strong knowledge of networking fundamentals and troubleshooting (Load Balancers, VPCs, Firewalls, DNS)
Experience implementing security best practices, including DDoS mitigation and incident response
Proficiency with reverse proxies and CDNs (Cloudflare, CloudFront, Akamai, etc.)
Hands-on experience with observability and monitoring stacks (ELK/EFK, Prometheus, Grafana, Datadog, New Relic)
Experience managing SQL/NoSQL databases and distributed systems (MySQL, Aurora, Redis, MongoDB, OpenSearch)
Experience implementing GitOps workflows (ArgoCD or Flux) and configuring scalable CI/CD pipelines using tools like GitHub Actions or GitLab Pipelines
Proficiency in scripting or programming with Bash, Python, Go, or JavaScript
Hands-on experience building and working with agentic AI development tools (i.e. GitHub Copilot, Claude Code, OpenCode, Cursor, Codex, etc.) to streamline day-to-day work (automation, troubleshooting, operational workflows, IaC, etc)