Drive the evolution of our Data Platform and lead migration to a new on-premise stack
Build, maintain, and migrate ETL/ELT pipelines from the legacy environment to the new stack
Work across the full toolchain: Python, SQL, Spark, Airflow, HDFS, Kafka, Trino, Iceberg, and dbt
Own Airflow as a production orchestration layer — DAGs, deployment, retries, sensors, pools, queues, callbacks/listeners, backfills, and reruns
Design and implement DataOps practices: monitoring, alerting, SLA/SLO tracking, runbooks, incident diagnostics, and postmortems
Set up observability across Airflow, Spark, HDFS, Kafka, and other platform components
Build custom listeners, exporters, checkers, and internal tooling for platform health diagnostics
Automate recurring team operations: DAG deployment, pipeline migrations, backfill/retry/recovery flows, and pre-release validation
Advance CI/CD and production-readiness standards for data workflows
Contribute to Data Governance at the engineering level — ownership, naming conventions, metadata, lineage, access patterns, auditability, and privacy-by-design
Work with Kubernetes at the application level: updating images, configuring deployments/jobs/cronjobs, migrating services, and managing configs, secrets, and env variables
Help the team cut down on manual ops, recurring failures, and operational noise