Strong understanding of cloud infrastructure and distributed computing principles
Experience with virtual machines, containerization, and managing compute resources
Experience building with IaC solutions, preferably Terraform
Knowledge of GPU clusters and techniques for optimizing ML workloads
Working knowledge of container orchestration systems like Kubernetes and job schedulers like SLURM
Familiarity with infrastructure components including networking, storage optimization, and resource management
Experience optimizing performance of diverse workloads in cloud environments
Strong programming skills, particularly in Python, and familiarity with the PyTorch ecosystem
Understanding of cloud infrastructure concepts and deployment patterns
Excellent written communication skills and ability to clearly express technical ideas in text
2+ years of experience in software development, cloud engineering, DevOps, or a similar technical role
Demonstrated experience with cloud technologies and infrastructure
Previous work with infrastructure-as-code, containerization, and cloud environments
Experience with MLflow, Apache Airflow, or Kubeflow
Familiarity with cloud ML platforms like AWS, GCP, Azure ML, or NVIDIA NGC
Experience managing hybrid cloud or on-prem GPU infrastructure
Background working with technology partners and integrating third-party solutions
Public presentation skills