You bring 10+ years of hands-on experience in Platform Engineering, Infrastructure Engineering, or building highly scalable distributed systems — with a strong foundation in software design, development, and algorithmic problem-solving
Data Platform Experience: Familiarity operating data platforms or data-intensive workloads, including distributed processing and streaming frameworks such as Spark, Airflow, Kafka, or Flink
Kubernetes & Containerization: Deep expertise in Kubernetes cluster design, day-to-day operations, and production troubleshooting across containerized service environments
CI/CD Systems: Proven track record building and operating robust CI/CD pipelines using tools such as Argo CD and GitHub Actions
Production Ownership: Demonstrated experience owning mission-critical systems with high availability requirements (≥99.99% uptime), including incident response, SLI/SLO/SLA definition, error budget management, and blameless postmortems
Distributed Systems: Hands-on experience developing large-scale distributed systems, databases, and backend APIs
Multi-Region Architecture: Practical experience designing and operating geo-replicated, active-active, multi-region systems — with a solid grasp of traffic routing, failover strategies, and data consistency tradeoffs
Observability: Strong experience building and owning full-stack observability solutions, including metrics, logging, and distributed tracing using tools such as Prometheus, Grafana, and OpenTelemetry
Infrastructure as Code: Proficiency with IaC tooling such as Helm, Terraform, or Pulumi, and experience with automated environment provisioning
Performance & Capacity: Strong command of system performance tuning, capacity planning, and resource optimization in distributed environments
Cloud-Native Security: Hands-on experience applying security best practices in cloud-native settings, including secrets management, network policies, and vulnerability scanning