Чем предстоит заниматься
Lead discovery and requirements-gathering sessions with business stakeholders; document source schemas and define analytics-ready data model specifications
Build, optimize and maintain end-to-end data pipelines ETL/ELT frameworks with reusable dbt macros, incremental models, and snapshot patterns to ensure fault-tolerant, scalable data delivery
Design and implement data models following medallion architecture patterns (bronze/silver/gold), creating foundational tables, star schemas, and reporting marts with defined SLAs and data quality contracts
Tune Spark jobs and pipeline components to meet SLA requirements by optimizing partitions, reducing shuffle, and implementing efficient transformation patterns
Implement automated data quality checks using frameworks like Great Expectations, establishing quality contracts across pipeline layers
Diagnose complex, multi-system pipeline issues—profiling Spark UI stages, tracing Kafka consumer lag, identifying dbt dependency failures—and design mitigations proactively
Create and maintain technical documentation, data model diagrams, and operational runbooks for production support