Z
AI Inference Engineer - Speech
Zoom
Seattle (WA), United States
🏢 Офис
Middle
Полная занятость
Продуктовая
США
Tech
Описание вакансии
Обязанности
•Accelerate deep learning inference on NVIDIA GPUs
•Design low latency high accuracy ASR systems
•Develop speech recognition services for products
•Improve inference latency throughput and memory usage
•Optimize ASR inference for production deployment
•Profile and debug ASR runtime bottlenecks
Технологии: ASR, BEAM Search, C#, C++, CUDA, CUDA Graphs, CUDA kernels, Deep learning, GPU clusters, Machine Learning, Mixed Precision, Model Optimization, NVIDIA GPUs, Profiling, PyTorch, Python, Real Time, Real-time Inference, Sequence to sequence, Shell, Speech Recognition, TensorFlow, TensorRT, Transformer