Advanced degree in a quantitative, computational, or life-science discipline
Experience working with biomedical, genomic, clinical, proteomic, imaging, laboratory, or other complex scientific data
Experience supporting research in life sciences, healthcare, diagnostics, or a similarly data-intensive and regulated scientific environment
Production experience with Spark or PySpark and distributed data processing
Depth in AWS services such as Athena, Glue, EMR, SageMaker, Lambda, Step Functions, Lake Formation, or related data and analytics technologies
Experience developing REST APIs or lightweight web applications used by non-engineering audiences
Experience designing automated validation frameworks, data contracts, reusable data-processing libraries, or researcher-facing workflow tools
Experience with metadata-management, data-catalog, or data-discovery platforms, such as the AWS Glue Data Catalog, OpenMetadata, or Unity Catalog
Experience with containerization, continuous integration and deployment, or infrastructure-as-code
Experience supporting machine-learning workflows or preparing data for model development and evaluation
Experience working with large files or multimodal datasets, such as sequencing outputs, digital pathology images, clinical records, or experimental measurements
Familiarity with governance considerations for research data, including access control, de-identification, and the handling of sensitive clinical information
Experience working within a data mesh, data product, or federated data-ownership model
Bachelor’s degree in computer science, data science, engineering, statistics, mathematics, bioinformatics, computational science, or another relevant quantitative discipline
Five or more years of relevant professional or applied research experience, or three or more years with an advanced degree in a relevant field
Advanced programming skills in Python
Strong SQL skills and experience working with structured and semi-structured data
Demonstrated track record of building reusable, maintainable software that others depend on, rather than one-time scripts or analyses
Experience designing and delivering several of the following: data pipelines, Python packages, APIs, analytical workflows, notebooks, or internal software tools
Substantial hands-on experience using AWS for data processing, analytics, scientific computing, or software development
Sufficient depth in AWS services and architecture to evaluate technical options, justify design recommendations, and define infrastructure requirements with DevOps or cloud-engineering partners
Experience conducting or supporting quantitative research, such as statistical analysis, machine learning, computational modeling, or another data-intensive research activity
Experience cleaning, integrating, standardizing, or validating data from multiple sources at meaningful scale
Fluency with software-development practices such as Git, automated testing, technical documentation, code review, and continuous integration
Demonstrated ability to investigate ambiguous problems, define an approach, and deliver a working solution with little guidance
Experience mentoring or providing technical guidance to other engineers, scientists, or analysts
Strong communication and collaboration skills, particularly when building alignment across scientific and technical disciplines