Hands-on experience in Data Science, including ownership of complex, multi-sprint projects for 3+ years
Bachelor's degree in Statistics, Data Science, Computer Science, Mathematics, or another quantitative field (Master's degree preferred)
Advanced proficiency in Python, including writing production-quality, well-tested, and well-documented code
Strong experience with SQL and PySpark, including processing datasets at billion-row scale
Hands-on experience with Databricks, including Workflows, Delta Lake, and job orchestration
Working knowledge of AWS or Google Cloud Platform (GCP)
Strong foundation in Machine Learning, including regression, classification, clustering, model evaluation, and experimental design
Experience with MLOps practices, including experiment tracking, Airflow-based pipeline orchestration, and reproducible model deployment
Familiarity with modern AI technologies, including Retrieval-Augmented Generation (RAG), LLM-based applications, vector databases, and semantic search
Strong written and verbal communication skills, including technical documentation, user story creation, and cross-functional collaboration
Level of English – from Intermediate+ and above
Knowledge graph construction, entity resolution, or semantic data modeling (RDF, OWL, SPARQL, or equivalent)
Experience with probabilistic record linkage, identity graph approaches, and embedding-based entity matching at scale
Experience with causal inference methods, including A/B testing, synthetic control, and uplift modeling
Experience with deduplication, data enrichment, or web-to-TV linkage problems
Background in media, ad tech, or measurement specifically: TV viewership (ACR/STB data), digital audience modeling, cross-platform measurement (linear + CTV/OTT), identity resolution in privacy-constrained environments
Familiarity with the measurement/identity vendor landscape: Nielsen, Comscore, LiveRamp, The Trade Desk