Описание вакансии
Обязанности
•Build dataset roadmap
•Collaborate with AI team on model training datasets
•Collaborate with scientists on data cost throughput and quality
•Extend GCP ingestion infrastructure
•Find new audio data sources
•Ingest audio data into ingestion pipeline
•Manage infrastructure with Terraform
•Operate ingestion pipeline infrastructure
Условия
•Asynchronous culture
•Flexible management approach
•Remote/distributed work
•Supportive team
Технологии: Bash, Data Ingestion, Data Processing, Docker, GCP, Large Scale Data, Large-scale, Large-scale Data Processing, Linux, Python, Terraform, Web Crawling