We are looking for a Data Engineer to help build and scale the data platform behind our search quality, ML pipelines, product analytics, and business operations
In this role, you will contribute across the full data lifecycle: ingesting data from production systems, designing and evolving our data warehouse, building batch and streaming pipelines, and making high-quality datasets available to researchers, engineers, analysts, and product teams across the company
The platform spans tens of terabytes and ingests data from tens of proprietary and third-party sources — including our search engine and its components, CRM, billing, identity, and product analytics across multi-region production environments. Around 100 internal users rely on it daily
You will work closely with engineers and stakeholders across the company, contribute to architectural and modeling decisions, and help improve the reliability, usability, and scalability of the data platform as it grows
Contribute to the design, development, and operation of Tavily's data platform — from real-time ingestion through data warehouse medallion layers to consumer-facing datasets and dashboards
Build and maintain reliable batch and streaming pipelines that ingest data from production services and external systems
Design and evolve scalable, analytics-ready data models in the data warehouse
Work closely with engineers across the company to ensure data produced by production systems is reliable, well-structured, and usable downstream
Improve observability across the data platform, including data quality checks, freshness monitoring, lineage, schema evolution, and cost controls
Partner with researchers, engineers, analysts, finance, and product managers to deliver trustworthy datasets for product, search quality, ML, and GTM analytics
Contribute to defining the objects, entities, and relationships that represent Tavily's search domain — including agent inputs, URLs, chunks, agent sessions, crawls, and the connections between them — and translate them into clean, queryable data models
Improve engineering practices around testing, documentation, deployment, and incident response
Investigate and resolve production data issues, including broken pipelines, corrupted datasets, schema changes, and large-scale backfills
Contribute to technical standards and best practices for data engineering across the company
Help maintain high standards of data quality, integrity, security, and governance across environments