Наши требования
Experience with large ML training pipelines or dataloading systems
Knowledge of columnar or custom data formats
Experience with systems like ClickHouse, Ray, Flink, Spark, or similar
Hands-on experience operating petabyte-scale datasets
Debugging and fixing performance bottlenecks in data-heavy systems
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records
Strong software engineering fundamentals
Experience building distributed systems or large-scale data pipelines
Comfort reasoning about performance, memory, I/O, and storage efficiency
Familiarity with batch and/or streaming processing systems
Experience with object storage systems and data format tradeoffs
Ownership mindset: design, build, operate, and iterate on systems end-to-end
Enjoy working closely with researchers and unblocking fast-moving projects