3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems
Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow)
Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores)
Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation)
Final offer amounts are determined by multiple factors including experience and expertise
Equity: In addition to the base salary, equity may be part of the total compensation package
ML compilers and framework internals: PyTorch internals, torch.compile, custom operators
Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism
Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving
Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis
Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads