Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue
Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog
Tune workers, partitioning, and shuffle behavior for cost and performance optimization
Model and load curated datasets into Snowflake with staging, transformation, and publishing layers
Implement automated data quality and validation frameworks, including schema/contract enforcement and null/uniqueness/referential checks
Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling
Configure and extend Glue Data Quality (DQDL) rules per requirements
Write clean, modular, testable Python with unit/integration tests and reusable libraries
Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager
Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring
Participate in code reviews, CI/CD automation, and documentation
Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and act as technical advisor within the workstream