Наши требования
Deep expertise building large-scale data analytics pipelines with Apache Spark
Comfortable using Python and AWS technologies for data engineering and ML
Strong communication skills to interact with technical counterparts across customer accounts to coordinate data integrations
Experience in the healthcare and life sciences domain
Familiarity with ML and statistical modeling methodologies and experimental design
Apache PySpark, Apache Iceberg, AWS EMR, Glue, S3
Python, PyTorch, sklearn, xgboost, catboost, mlflow