Компания скрыта (Fintech)·Worldwide·3 дн назад
Data Engineer
5 000 – 7 000 $
на 152% выше медианы рынка
≈ 419,9 тыс.–587,8 тыс. ₽
🌍 УдалённоSeniorПолная занятостьАнглийский B2
59
Хорошие условия
Навык редкий и дорогой (Kubernetes), работа расписана подробно. Но верх вилки на 66% ниже медианы грейда.
Вакансия на английскомПереведёт заголовок и описание вакансии на русский
Наша компания
The Data Engineering team is responsible for designing, building, and maintaining the Market Data Platform — a lakehouse infrastructure spanning the full path from raw exchange feeds to reliable, petabyte-scale data for research, backtesting, and real-time trading.
О роли
Capture & Ingestion. Own the full capture path from wire to lake: decode and normalize raw exchange feeds (PCAP, multicast UDP, ITCH, FIX) and vendor sources (OneTick, Refinitiv, Bloomberg, ICE) into a unified canonical model with nanosecond timestamps. Build batch and streaming pipelines (Airflow, Spark, dbt) for tick and reference data. Own L2/L3 order-book reconstruction with gap handling. Provide Python and Rust producer SDKs for internal feed handlers
Чем предстоит заниматься
Storage & Modeling — Apache Iceberg. Own the Iceberg-over-S3 lakehouse: design partitioning, sort orders, and row-group layouts for fast scans; manage schema evolution, snapshots, time travel, compaction, and TTL. Maintain reference data as slowly changing tables with point-in-time correctness for backtests. Drive storage cost optimization through compaction, tiering, and snapshot expiry
Tooling & Libraries. Build libraries for schema management, data contracts, validation, and lineage on top of the Iceberg catalog. Develop shared access services (Spark + Polars) so research, backtesting, and trading share a single normalized data layer, including gap detection and PCAP-vs-lake reconciliation
Reliability & Observability. Embed monitoring, alerting, SLAs/SLOs, and CI/CD across capture and pipeline layers on Kubernetes (EKS). Own data-quality dashboards and incident runbooks for the capture fleet
Collaboration. Partner with Quant Research, Data Science, Backend, and DevOps to translate requirements into platform capabilities and champion market-data engineering best practices
Наши требования
5+ years of experience building production-grade data systems, with deep knowledge and understanding of the problems that come with running data lakes/lakehouses at scale: data layout and partitioning strategy, ingestion and backfills, small-file and metadata growth, consistency and late-arriving data, query performance, storage cost, and operational failure modes
Hands-on experience with Apache Iceberg (or comparable table formats such as Delta/Hudi): partitioning, schema evolution, snapshots, compaction, and catalog operations; familiarity with Apache Arrow for zero-copy, columnar in-memory interchange
Expert-level Python (incl. Polars and/or PySpark)
Modern orchestration (Airflow) and distributed processing/query engines (Apache Spark, StarRocks)
Advanced SQL: complex aggregations, window functions, query optimization, partition pruning
Solid fundamentals in Linux, containerization (Docker, Kubernetes/EKS), and cloud object storage (AWS S3)
DevOps & observability: CI/CD, infrastructure-as-code (Terraform), GitOps (ArgoCD), and metrics/dashboards/alerting (Grafana, Prometheus)
Strong grasp of structured + unstructured/binary data and storage optimization: partitioning, compression, cost management
English fluency (B2+) for documentation and collaboration in an international team
Experience with market data and/or network packet capture: decoding PCAP, exchange feed protocols (ITCH, FIX/FAST, multicast UDP), order-book reconstruction, and time-series at scale. Willingness to learn this domain is expected
Experience normalizing market data from multiple vendors (OneTick, Refinitiv/Reuters, Bloomberg, ICE) into a unified schema and symbology
Rust, relevant for high-performance capture/decoding
Мы предлагаем
Fully remote setup with a daily overlap window from 12:00 to 18:00 GMT+3
Reimbursement for health insurance, sports, and personal development
A genuinely flat structure with real autonomy over how you work
Room to grow the role as the team scales
