5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering
Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana
Hands-on experience with Langfuse, Grafana, and Prometheus
Experience with Terraform and CI/CD
Python skills for instrumentation, exporters, and automation
Familiarity with ML workloads and AI-specific metrics
Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs
Intermediate or higher English
PromQL and Kusto Query Language
OpenTelemetry, including GenAI semantic conventions
LLM evaluation frameworks
AI cost dashboards and FinOps
Alerting, on-call, and incident management tooling
AKS and Kubernetes observability