3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems
Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events
Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers
Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines
Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations
A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams
Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations
English Level: B2+ (Upper-Intermediate) or higher
Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
Experience with Datadog or similar enterprise observability platforms
Background in evangelizing best practices and setting standards across engineering teams
Exposure to programmatic advertising or adtech platforms