Чем предстоит заниматься
Build data pipelines and notebooks
Build reliable restartable ingestion processes
Conduct code reviews and pair programming
Develop SQL Python PySpark Spark SQL transformations
Document solutions and naming conventions
Implement lakehouse and warehouse medallion architecture
Manage deployments and Git workflows
Model data into dimensional star schemas
Monitor optimize and troubleshoot data pipelines
Set up CI/CD pipelines and release processes
Unlock source systems