We're looking for a Data Engineer to build and maintain the data pipelines that feed Sesame's AI models. You'll collaborate directly with machine learning engineers and researchers — your job is to make sure they have the right data, in the right shape, at the right time to train, evaluate, and ship models
Sesame's data is rich and complex: conversations, voice, sensor signals, and product telemetry. You'll design the systems that take raw, unstructured, multimodal data and turn it into clean, versioned, well-documented datasets that ML teams can trust and build on confidently
This is a deeply technical, infrastructure-focused role — closer to ML engineering than traditional data analytics. You'll be deeply embedded with ML teams, understanding their workflows and building infrastructure that accelerates the full model development lifecycle — from data collection and labeling through training and evaluation
Design and build production data pipelines that prepare conversational, voice, and multimodal data for model training and evaluation
Partner directly with ML engineers to understand data requirements for new models and experiments, and deliver datasets that meet those needs
Build and maintain infrastructure for dataset versioning, lineage tracking, and reproducibility — so any training run can be traced back to its exact data
Develop data quality frameworks that catch issues before they become model quality issues: schema validation, drift detection, and coverage monitoring
Optimise large-scale data processing for cost and performance across Sesame's cloud infrastructure
Build tooling that makes it easy for ML engineers and researchers to discover, explore, and request data independently
Define and enforce data governance and privacy standards, particularly around sensitive conversational and voice data
Contribute to architecture decisions around Sesame's broader data platform as the team and data volume grow