Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC)
Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code
Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production
Distributed systems: High-throughput control planes, microservices, or multi-region setups
Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture
Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate
Breadth over comfort: Willingness to dig into unfamiliar parts of the stack when a problem crosses boundaries
Education: Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent practical experience
Observability tooling: Prometheus, Grafana, OpenTelemetry, and alerting people actually act on
GPU and ML infrastructure exposure: GPUs, inference serving, or distributed training
AI-assisted operations: Building agents or LLM-based tooling for investigation, triage, or automation
Open source background: Contributions to infrastructure, systems, or ML serving projects
Startup agility: Comfortable where pragmatism and teamwork matter more than process