Чем предстоит заниматься
Support a team of data engineers to build pipelines used by MLOps and ML Engineers on ML modeling teams
Develop, optimize, and maintain data transformation pipelines in Databricks using PySpark
Work with data stored in ADLS Gen2 and SAP HANA Data Lake, primarily in Delta/Parquet format
Implement and maintain data quality checks, including schema validation, deduplication, enrichment, and tagging
Communicate with stakeholders to understand business processes and model input data
Tune performance for large-scale datasets to ensure efficient processing