You’ve run multi-cluster Kubernetes on EKS in production — backend, networking, persistence, monitoring, models, and kafka clusters per region — and you’ve used Cluster API or similar for programmatic cluster creation
You’ve operated a stateful data plane (Postgres, Redis, Kafka, Temporal, etcd, ClickHouse) at scale — you’ve sharded it, migrated data between instances, and lived with the consequences
You’re fluent in Envoy and Cilium/eBPF. You’re comfortable debugging Envoy response flags, conntrack drops, and cross-zone LB behavior. VPC/NAT/Cloudflare alone isn’t enough
You’ve run multi-region Kafka on MSK in production — not just Kafka. You’ve dealt with regional topic naming, MSK Pulumi drift, and compliance constraints
You write Go for control-plane services. Vapi’s cluster-manager, traffic-control-plane, and environment-manager are all Go, and you’re comfortable owning code in that stack
Bonus: SIP / RTP / telephony background. The Nov 7 SIP gateway SPOF is still unsolved, and a telephony-savvy infra hire unblocks that roadmap item
Bonus: cell-based / shard architecture experience — Shopify pods, AWS cell-based reference arch, Slack shards, or equivalent. Microservices experience alone isn’t the same
You likely come from one of: a company that ran cell-based in prod (Shopify, AWS service teams, Slack); a distributed systems shop (Cockroach, MongoDB, Confluent, Temporal, Redpanda, ClickHouse Inc.); a voice/video/CPaaS company (Twilio, Plivo, Bandwidth, Vonage, LiveKit, Daily.co Dialpad); an Envoy/service-mesh org (Lyft, Stripe, Airbnb, Pinterest, Isovalent/Cilium); or a streaming-infra team (Confluent, Uber, LinkedIn, Datadog) that ran MSK/Kafka multi-region