Staff/Principal DevOps Engineer, AI Inference
Core
Design, implement, and optimize infrastructure for serving machine learning models at scale, bridging platform engineering, SRE, and ML infrastructure to power low-latency, high-throughput inference across GPU clusters.
Role type
Staff/Principal DevOps Engineer (AI Inference)
Builds
GPU/accelerator infrastructure on Kubernetes, model serving platforms, intelligent request routing, autoscaling systems, production-grade deployment pipelines, and AWS cloud infrastructure for ML.
Domain
AI/ML Infrastructure, Cloud Computing, GPU Acceleration
Deliverable
production ML models
Required skills
Kubernetes (GPU scheduling, resource quotas), AWS (EKS, EC2 P-series/Inf/Trn, Terraform, Helm), Model serving frameworks (vLLM, Triton, TGI), Python, Networking (NCCL, VPC, load balancing), Infrastructure as Code, Observability
Preferred skills
LLM inference optimization (continuous batching, speculative decoding, quantization), Multi-accelerator experience (NVIDIA, AWS Inferentia/Trainium, AMD), Rust or Go, Chaos engineering, ML supply chain security
Technologies
Kubernetes, AWS (EKS, EC2, S3, EFA), Terraform, Helm, vLLM, Triton Inference Server, TGI, Python, Rust, Go, NCCL
Responsibilities
Schedule and manage multi-tenant GPU workloads on Kubernetes; Deploy and optimize model serving platforms with batching and caching; Implement autoscaling for heterogeneous accelerator fleets; Build production deployment pipelines with canary rollouts and versioning; Manage AWS GPU infrastructure and networking; Optimize cost and capacity for inference workloads.
Seniority
Staff/Principal, hands-on IC with strategic impact