8-10 years of experience in Infrastructure Engineering or similar roles (DevOps, Systems Engineering, Site Reliability Engineering)
Strong programming skills in languages like Python or Go
You write high-quality, well-tested code
Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture
Experience with container orchestration platforms (Kubernetes) and cloud-native technologies
Proven track record of implementing and maintaining monitoring/observability solutions, with strong skills in debugging and performance tuning
Strong incident management skills with experience leading incident response and demonstrated critical thinking under pressure
Experience with infrastructure as code (e.g., Terraform) and configuration management tools
Excellent written and verbal communication skills, with an ability to explain technical concepts clearly and simply and a bias toward open, transparent cultural practices
Strong interpersonal skills, with experience working with engineers from junior to principal levels
A willingness to dive into understanding, debugging, and improving any layer of the stack
You're passionate about making software creation accessible and empowering the next generation of builders
Deep experience with Google Cloud Platform (GCP) services and tools
Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.)
Experience designing and building reliable systems capable of handling high throughput and low latency
Experience with Go and Terraform
Familiarity with working in rapid-growth environments
Experience writing company-facing blog posts and training materials
This is a full-time role that can be held from our Foster City, CA office. The role has an in-office requirement of Monday, Wednesday, and Friday