5+ years of experience operating Linux systems in production or HPC environments, with hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS, or similar)
Hands-on experience operating Software-Defined Storage (SDS) platforms at scale, including integrating with their management and data-plane APIs
Strong incident-response instincts: comfortable owning a production storage incident end to end, from first alert through root cause to postmortem
Working experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic — including building dashboards and alert/pager routing for multiple audiences
Working experience with Kubernetes (GitOps tooling such as ArgoCD, Helm/Kustomize) and hands-on troubleshooting
Working experience with CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming in Python or Go
Working experience with Infrastructure as Code (Terraform, Ansible)
Solid understanding of core storage protocols across file (NFS, SMB), object (S3), block (NVMe-oF/TCP), and structured (vector DB, SQL) storage
Experience with cutting-edge software-defined storage solutions such as VAST or Weka
Enterprise storage expertise in solutions such as NetApp, Dell PowerScale, GPFS, or Lustre
Experience writing or operating Kubernetes CSI drivers
Experience with SR-IOV and virtualization (KVM/QEMU)
Experience with GPUDirect Storage, RDMA, InfiniBand, or RoCE networking
Familiarity with NIC-level diagnostics (ethtool, mlxlink) and fleet-wide operations tooling (clush or similar)
Contributions to open-source storage projects
Salary Range Information