Software Engineer, Model Inference
Core
Optimizing the world's largest AI models for high-volume, low-latency, and high-availability production and research environments.
Role type
Senior IC machine-learning inference engineer
Builds
Production inference stack, tools for visibility into bottlenecks, and optimized code/fleet for Azure VMs
Domain
Artificial Intelligence / High-Performance Computing
Deliverable
production ML models
Required skills
Modern ML architectures, PyTorch, NVidia GPUs, CUDA, NCCL, HPC technologies (InfiniBand, MPI, NVLink), distributed systems architecture, debugging production systems, system refactoring at scale
Preferred skills
Performance-critical distributed systems experience, self-directed problem solving
Technologies
Azure VMs, PyTorch, NVidia GPUs, CUDA, NCCL, InfiniBand, MPI, NVLink
Responsibilities
Collaborate with researchers to bring latest technologies into production, introduce new techniques/tools/architecture to improve inference performance/latency/throughput/efficiency, build tools for bottleneck visibility and implement solutions, optimize code and hardware fleet utilization
Seniority
Senior, hands-on IC