Design, build, andmaintainscalable, resilient data and ML pipelines, infrastructure, and workflows using tools such as Terraform, GitHub Actions,ArgoCD, Helm, and others
Automate infrastructure provisioning and configuration management using cloud-native services (preferably AWS) with tools like Terraform, CloudFormation
Design, containerize, and manage Kubernetes (EKS) clusters and/or ECS environments in AWS. Collaborate with development teams tooptimizeperformance, deployment, and cost
Partner with DevOps and SRE teams to ensure high availability, observability, scalability, and security of the data and ML infrastructure
Work closely with Data Scientists and ML Engineers to operationalize machine learning models, including building CI/CD pipelines for model training, validation, and deployment
Implement observability for data pipelines and ML services using tools like Prometheus, Grafana, Datadog, or similar
Develop andmaintainautomated pipelines for model retraining, monitoring drift, and versioning in production
Support experimentation and prototyping in areas such as Machine Learning and Generative AI, transitioning successful prototypes into production systems
Ensure cloud infrastructure is secure, compliant, and cost-efficient, following best practices in governance, identity, and access management