As a Senior Data Engineer, you will play a crucial role in building and maintaining the foundation of our data ecosystem. You’ll work alongside data engineers, analysts, and product teams to create robust, scalable, and high-performance data pipelines and models. Your work will directly impact how we deliver insights, power product features, and enable data-driven decision-making across the company
This role is perfect for someone who combines deep technical skills with a proactive mindset and thrives on solving complex data challenges in a collaborative environment
Experience with additional AWS services: EMR, EKS, Athena, EC2
Hands-on knowledge of alternative data warehouses like Snowflake or others
Experience with PySpark for big data processing
Familiarity with event data collection tools (Snowplow, Rudderstack, etc.)
Interest in or exposure to customer data platforms (CDPs) and real-time data workflows
Candidate journey: ⭕️ Recruiter call ➔ ⭕️ Technical call with the hiring manager ➔ ⭕️ Meet the future stakeholders
Check out some of our products
4+ years of experience in data engineering or backend development, with a strong focus on building production-grade data pipelines
2-3+ years of experience working with AWS services (Administration of Redshift is a must)
Solid experience working with AWS services (Spectrum, S3, RDS, Glue, Lambda, Kinesis, SQS)
Proficient in Python and SQL for data transformation and automation
Experience with dbt for data modeling and transformation
Good understanding of streaming architectures and micro-batching for real-time data needs
Experience with CI/CD pipelines for data workflows (preferably GitLab CI)
Familiarity with event schema validation tools/ solutions (Snowplow, Schema Registry)
Excellent communication and collaboration skills
Strong problem-solving skills—able to dig into data issues, propose solutions, and deliver clean, reliable outcomes
A growth mindset—enthusiastic about learning new tools, sharing knowledge, and improving team practices
Cloud: AWS (Redshift, Spectrum, S3, RDS, Lambda, Kinesis, SQS, Glue, MWAA)
Languages: Python, SQL
Orchestration: Airflow (MWAA)
Modeling: dbt
CI/CD: GitLab CI (including GitLab administration)
Monitoring: Datadog, Grafana, Graylog
Event validation process: Iglu schema registry
APIs & Integrations: REST, OAuth, webhook ingestion
Infra-as-code (optional): Terraform