8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering)
Strong programming skills in languages like Python or Go. You write high-quality, well-tested code
Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture
Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies
Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing)
Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure
Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools
Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices
Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels
A willingness to dive into understanding, debugging, and improving any layer of the stack
You're passionate about making software creation accessible and empowering the next generation of builders
Deep experience with Google Cloud Platform (GCP) services and tools
Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)
Experience designing and building reliable systems capable of handling high throughput and low latency
Significant experience with Go and Terraform
Familiarity with working in rapid-growth, startup environments
Experience writing company-facing blog posts and training materials