Own the Technical Blueprint: Personally architect the infrastructure solutions for our most strategic M&E partnerships , studio-scale content production pipelines, agency data consolidation plays, generative AI model deployments. These architectures must be engineered to survive real-world scale, not just pass a POC
The Physics to P&L Narrative: Fluently demonstrate to executive stakeholders how infrastructure decisions , data lake locality, storage tiering, inference optimization, directly impact their business model and operability
Shape the M&E Roadmap: Use forensic evidence from the field to prioritize and justify the M&E vertical roadmap. You will work directly with Nebius’s global Head of Product and Head of Engineering to translate partner and customer needs into product direction
Lead the M&E Product Summit: Chair a quarterly summit with Core Engineering leadership, using field evidence to drive roadmap decisions and maintain vertical momentum
Mastery of the Stack: Expert-level, production-grade knowledge of GPU architectures (H100, L40s), Kubernetes orchestration including Soperator, high-performance and parallel file systems (e.g., Lustre, WEKA), data lake architecture, and networking constraints (InfiniBand/Ethernet)
Inference Optimization: You understand the nuances of model serving , batch sizes, quantization, KV caching, latency tradeoffs , and can architect solutions for both massive throughput and real-time (sub-50ms) demands
Domain Context: Media & Entertainment
Industry Fluency: You have operated within the M&E ecosystem and can speak to the infrastructure implications across the following sub-sectors
Gaming
Multimodal Generative Models
AdTech
VFX / Content Production
Non-negotiable: Candidates without deep, production-grade working knowledge of GPU infrastructure and inference-based solutioning will not be considered. This is a hard requirement, not a preference
Mastery of the Stack: You must be an expert in the physics of AI infrastructure. This includes deep, production-grade knowledge of GPU architectures (H100 vs. L40s), Kubernetes orchestration (K8s), high-performance storage parallel file systems (e.g., Lustre, WEKA), and networking constraints (InfiniBand/Ethernet)
Inference Optimization: You understand the nuances of model serving—batch sizes, quantization, KV caching, and latency tradeoffs—and can architect solutions for both massive throughput and real-time (sub-50ms) demands
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Pay Transparency
Base Compensation Range
$200,000—$245,000 USD