Direct experience with end-to-end LLM fine-tuning, algorithm selection, pipeline design, or distributed training setups (3D parallelism, Megatron-LM)
Prior experience with product launches, leading GTM initiatives, or publishing technical whitepapers/benchmarks
Experience integrating RESTful APIs, gRPC, and service-oriented cloud architectures
Salary Range Information
Have a proven track record deploying, benchmarking, and optimizing workloads on NVIDIA GPU architectures (e.g., HGX platforms, NVLink) using deep learning frameworks (PyTorch, NeMo) and inference engines (vLLM, TensorRT-LLM)
Have 8+ years of experience designing, deploying, and scaling enterprise cloud infrastructure
Have 4+ years in a Solution Architect, Solution Engineer, or technical customer-facing capacity supporting complex cloud environments
Have 3+ years of hands-on experience architecting and deploying cloud-based AI/ML workloads
Have strong experience with modern infrastructure orchestration tools such as Kubernetes, Docker, SLURM, Terraform, and Ansible
Have deep knowledge of cloud networking concepts, including high-speed interconnects (InfiniBand, RoCE), distributed file systems (NFS, NVMe-oF, Weka, VAST), security, and cost optimization
Have experience coding in Python, Go, C/C++ (CUDA) or similar programming language
Have experience partnering with Account Executives to close complex cloud deals, present technical architectures to C-level stakeholders (CTOs, VP of Eng), and drive customer alignment
Have demonstrated impact at an organizational/multi-departmental level and are effective mentoring junior SEs or architects
Thrive in dynamic settings and embrace radical ownership of initiatives and outcomes