Design and evolve Kubernetes control plane architecture across regions
Define and implement multi-tenant cluster models, including shared control planes, virtual cluster approaches (e.g., vcluster, Kamaji)
Drive transition from standalone clusters to regionally managed platform models
Define standards for isolation boundaries, resource segmentation, policy enforcement
Own the reliability and behavior of Kubernetes platforms in production
Participate in on-call rotation and lead incident response
Diagnose and resolve control plane instability, API server saturation, scheduling and resource contention issues
Ensure consistent lifecycle management across clusters - provisioning, upgrades, scaling
Design and implement strategies for regional scaling, multi-data center cluster deployments
Ensure consistent behavior and reliability across environments
Define cluster topology and failure domain strategies
Design ingress and egress architectures at cluster level and regional level
Troubleshoot and optimize pod-to-pod networking, north-south traffic flows, CNI behavior (Cilium preferred)
Collaborate with network engineering on high-performance networking integration
Improve observability across control plane components, cluster health and performance
Define and implement resilience strategies aligned with platform goals
Lead root cause analysis for production incidents