Чем предстоит заниматься
Design, implement and maintain end-to-end ML pipelines on Databricks
Build workflows for data ingestion, preprocessing, feature engineering, training and inference
Leverage PySpark, Spark ML and Databricks notebooks/jobs
Manage model versioning, experiment tracking and reproducibility using MLflow
Package and deploy models for batch and real-time inference
Monitor model performance, drift and retraining cycles
Develop scalable ETL/ELT pipelines using Databricks Delta Lake
Optimize data storage and access patterns through partitioning, Z-ordering and caching
Integrate with data sources such as Azure Data Lake, S3, APIs and databases
Implement CI/CD pipelines for ML workflows using Azure DevOps, GitHub Actions and Databricks Repos and Jobs API
Configure clusters, autoscaling and cost optimization while applying Infrastructure as Code with Terraform, ARM and Bicep
Implement logging, alerting and observability to ensure high availability and fault tolerance of ML systems