6+ years as a DevOps Engineer / SRE (or very close responsibilities)
Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery. You can showcase initiatives that were yours, not just tasks you completed
Confident Linux skills (we use Ubuntu)
Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup
Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery
Containers: Docker, image building, registries
Ansible
Git
Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work)
Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems
Experience mentoring less experienced engineers
Ownership and attention to detail. Downtime is expensive: during busy events 10 minutes of downtime can cost us around $500k
We understand it’s impossible to be an expert in everything, but it’s important to have solid hands-on experience in two or more of the areas below
VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure
Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (Fluent Bit or similar), retention, sharding, performance
Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets
Container registries and artifact management: Harbor, Nexus, base images, image policies
Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting
Grafana: dashboards as code, alerting, performance at scale
Great if you’ve worked with any of the following
Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits. Building platform around these tools to improve Quality of Life for Analytics
Remote development environments and AI agent execution environments. E.g. Coder/Telepresence
Bare-metal Kubernetes: provisioning, networking, scaling
Flux and GitOps
Terraform
Sentry on-premise: operating self-hosted error tracking
ClickHouse, MongoDB