CareerPlanSign in

Staff/Principal DevOps Engineer, AI Inference

Alewife, Cambridge, MA💼 Full-time💰 $192,000–$192,000🗓 2026-07-28 → 2026-09-26

Core

Design, implement, and optimize infrastructure for serving machine learning models at scale, bridging platform engineering, SRE, and ML infrastructure to power low-latency, high-throughput inference across GPU clusters.

Role type

Staff/Principal DevOps Engineer (AI Inference)

Builds

GPU/accelerator infrastructure on Kubernetes, model serving platforms, intelligent request routing, autoscaling systems, production-grade deployment pipelines, and AWS cloud infrastructure for ML.

Domain

AI/ML Infrastructure, Cloud Computing, GPU Acceleration

Deliverable

production ML models

Required skills

Kubernetes (GPU scheduling, resource quotas), AWS (EKS, EC2 P-series/Inf/Trn, Terraform, Helm), Model serving frameworks (vLLM, Triton, TGI), Python, Networking (NCCL, VPC, load balancing), Infrastructure as Code, Observability

Preferred skills

LLM inference optimization (continuous batching, speculative decoding, quantization), Multi-accelerator experience (NVIDIA, AWS Inferentia/Trainium, AMD), Rust or Go, Chaos engineering, ML supply chain security

Technologies

Kubernetes, AWS (EKS, EC2, S3, EFA), Terraform, Helm, vLLM, Triton Inference Server, TGI, Python, Rust, Go, NCCL

Responsibilities

Schedule and manage multi-tenant GPU workloads on Kubernetes; Deploy and optimize model serving platforms with batching and caching; Implement autoscaling for heterogeneous accelerator fleets; Build production deployment pipelines with canary rollouts and versioning; Manage AWS GPU infrastructure and networking; Optimize cost and capacity for inference workloads.

Seniority

Staff/Principal, hands-on IC with strategic impact

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.