Чем предстоит заниматься
Build data pipelines for AI/ML
Collaborate with data science and AI platform teams
Contribute to model validation and improvements
Design data ingestion and retrieval ready datasets
Design distributed systems for data consistency and throughput
Develop ETL and ELT workflows
Drive delivery of AI ML use cases
Implement data quality checks
Maintain Delta Lake and Iceberg tables
Manage lakehouse data partitioning and file formats
Monitor production data pipelines
Participate in governance and risk management
Process unstructured data for AI consumption
Support data warehouse and data lake environments
Troubleshoot and trace pipeline failures
Tune Apache Spark jobs