Чем предстоит заниматься
Design, develop, and maintain robust and scalable data pipelines to support various data-driven initiatives
Collaborate with cross-functional teams to understand data requirements and implement solutions that meet business needs
Work with large datasets, ensuring data quality, integrity, and reliability throughout the entire data lifecycle
Leverage Apache Spark within the Databricks environment to process large volumes of data efficiently
Utilize Python, SQL, Spark, and other technologies to optimize and enhance data processing capabilities
Implementation experience on containerization and orchestration tools (e.g., Docker, Kubernetes)
Implement and optimize ETL processes for efficient data extraction, transformation, and loading
Collaborate with Data Scientists and Analysts to provide them with the necessary data infrastructure and support their analytical needs
Stay current with industry trends and best practices to continuously improve data engineering processes