Anomaly Detection & Heuristics Expertise: Deep experience building anomaly detection systems, heuristics-based rule engines, or ML/RL systems for infrastructure or data-intensive domains
Threshold & Signal Calibration: Demonstrated ability to reason about precision/recall trade-offs and build feedback loops that keep detection systems accurate over time
Distributed Systems Fundamentals: Strong foundations in the building blocks of reliable, scalable backend systems—you can hold your own in any system design conversation
Full Software Engineering Craft: 5+ years shipping production software; experience with modern compiled or systems languages (Go, Rust, C++, Java, or similar)
Data & Observability Fluency: Comfortable with time-series data, telemetry pipelines, and observability primitives—you understand how raw metrics become actionable insights
Communication: You can explain detection logic, trade-offs, and system behavior clearly to both engineers and non-technical partners
Force Multiplier Mindset: You make the team better—through mentorship, clear technical vision, and a genuine investment in the people around you
Experience with GPU profiling tools (Nsight, NCCL Inspector) or hardware-level infrastructure diagnostics
Background in observability platforms or products
Experience with reinforcement learning applied to operational or infrastructure problems
Familiarity with large-scale fleet management or cloud infrastructure
Passion for building team culture and engineering quality of life