You’ve run incident command and postmortem discipline at scale on a real oncall rotation
You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog
You’ve done capacity planning and load testing for production systems with real users
You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown
You know backpressure and autoscaling patterns — KEDA, custom metrics scaling
You ship code, not just scripts. You can build platform services in Go or TypeScript (matches Vapi’s cluster-manager, database-health, wscaler, incidentManager)
Real-time / latency-sensitive product background where degraded means a dropped call, not a slow dashboard
Languages: Go and TypeScript (you ship code, not just scripts), Bash
Observability: Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry
Orchestration: Kubernetes on EKS — production ops (HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown, pod crash diagnosis)
Autoscaling and backpressure: KEDA, custom metrics scaling (matches Vapi’s wscaler and workerpool-cron-scaler)
Load testing: script-based load testing, provider rate-limit auditing, per-org concurrency auditing
Vapi services you’ll touch or build: cluster-manager, database-health, wscaler, incidentManager