4–5+ years of data engineering or general software development experience in a commercial setting
Strong hands-on proficiency in both Java and Python for production systems
Demonstrated experience designing and implementing complex ETL/ELT processes from concept to production
Proficiency in SQL and data exploration across large, complex datasets
Experience with big data technologies (Hadoop, Hive, BigQuery, Snowflake) and working with terabyte-scale or larger datasets
Solid foundation in data structures, algorithms, and object-oriented design
Professional experience building REST services and working with event queue systems
Familiarity with Linux and cloud infrastructure design on AWS or equivalent cloud providers
Systems performance and tuning experience, with an understanding of how architecture impacts scalability
Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience
Strong analytical skills; desire to write clean, correct, and efficient code
Self-paced, organized, and detail-oriented with a strong sense of ownership, urgency, and pride in your work
Ability to break down complex problems into simple, pragmatic solutions
Strong interpersonal skills, intense curiosity, and enthusiasm for solving difficult problems
Willingness and ability to take on new technologies as our stack evolves
Experience with stream processing frameworks such as Flink or Spark Streaming
Experience with task orchestration tools such as Apache Airflow or similar systems
Familiarity with additional technologies: GraphQL, React, HTML5, JavaScript, CSS, Postgres, Gradle, BERT
Experience working with and designing infrastructures for large-scale data processing (Hive, Snowflake, NoSQL databases)
Experience with data governance practices and tooling
Exposure to and/or interest in machine learning, data science, or GenAI to help solve engineering challenges innovatively
Experience developing scalable code for high-volume, low-latency systems