Design, build, and maintain a global, distributed, and resilient cloud infrastructure
Collaborate with infrastructure and product engineering teams to plan and deliver complex platform initiatives
Participate in architecture reviews, incident response, and performance analysis to ensure system reliability
Manage and provision AWS infrastructure using Terraform and Kubernetes
Write and maintain Kubernetes manifests and deployment configurations for critical workloads, including pod security contexts, anti-affinity rules, network policies, autoscaling, and health probes
Drive Production Readiness Reviews (PRR) for all new services, covering security, HA, performance, and observability gates
Design and operate multi-layer workload isolation using Linux kernel primitives: namespaces (pid, net, mnt, user, uts, ipc), cgroups, seccomp profiles, and capabilities, as the baseline security boundary
Evaluate and operate gVisor and Firecracker for workloads requiring hard tenant boundaries and near-native performance
Design and maintain Grafana dashboards
Manage Prometheus and VictoriaMetrics pipelines; define and tune P1/P2/P3 alert thresholds with runbooks
Contribute to and extend the internal k6-based load testing framework (load-testing-framework / library/k6/webhooks)
Design load scenarios using constant-arrival-rate profiles; instrument custom metrics (job_succeeded_count, job_failed_count) tagged by testid for Grafana correlation; stream test metrics to Prometheus via remote write; generate and publish HTML reports to file storage after each run