Experience with dbt and a close working relationship with analytics engineering teams
Experience with lakehouse architectures and open table formats (Iceberg, Delta Lake) or query engines like Trino
Experience operating multi-tenant platforms with strict security, compliance, or data residency requirements
Exposure to data infrastructure for AI products
Prior experience as an early or founding data platform hire at a fast-growing company
5+ years building and operating production data infrastructure, with ownership of systems other teams depend on
Deep experience with cloud data warehouses — Snowflake strongly preferred (BigQuery, Databricks, or Redshift experience transfers well) — including performance tuning and cost management
Hands-on experience building CDC and streaming pipelines with technologies like Kafka, Debezium, Flink, or Spark Streaming
Experience with managed ingestion tooling (Fivetran, Airbyte, or similar) and clear judgment about when to buy the connector and when to build it
Strong fluency with workflow orchestration — Temporal, Airflow, Dagster, or similar — operated at scale, not just configured
Strong programming skills in Python and advanced SQL
Experience building frameworks or internal tooling that other engineers use, and the product instinct to know when an abstraction is helping versus getting in the way
Practical experience with data quality, observability, and lineage tooling, and with schema evolution in systems that can't afford downtime
Working knowledge of data governance in a regulated environment: PII classification, masking, access control, retention, and data residency
Familiarity with cloud data services (Azure, AWS, GCP), Kubernetes, and infrastructure-as-code (Terraform, Pulumi)
Comfort operating in ambiguity and defining scope where none exists