CareerPlanGet AI match score →

Software Engineer, Model Inference

San Francisco💼 Full-time🗓 2025-02-06 → 2026-07-31

Core

Optimizing the world's largest AI models for high-volume, low-latency, and high-availability production and research environments.

Role type

Senior IC machine-learning inference engineer

Builds

Production inference stack, tools for visibility into bottlenecks, and optimized code/fleet for Azure VMs

Domain

Artificial Intelligence / High-Performance Computing

Deliverable

production ML models

Required skills

Modern ML architectures, PyTorch, NVidia GPUs, CUDA, NCCL, HPC technologies (InfiniBand, MPI, NVLink), distributed systems architecture, debugging production systems, system refactoring at scale

Preferred skills

Performance-critical distributed systems experience, self-directed problem solving

Technologies

Azure VMs, PyTorch, NVidia GPUs, CUDA, NCCL, InfiniBand, MPI, NVLink

Responsibilities

Collaborate with researchers to bring latest technologies into production, introduce new techniques/tools/architecture to improve inference performance/latency/throughput/efficiency, build tools for bottleneck visibility and implement solutions, optimize code and hardware fleet utilization

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗