CareerPlanGet AI match score →

Senior AI Engineer, SRE & LLM Infrastructure

💼 Full-time🗓 2026-05-29 → 2026-05-31

Core

Ensuring the reliability, scalability, and cost-efficiency of large language model (LLM) serving in production environments.

Role type

Senior SRE / Production Engineer specializing in AI Infrastructure

Builds

High-availability LLM serving stacks on GPU-accelerated infrastructure

Domain

Artificial Intelligence / Large Language Models / Cloud Infrastructure

Deliverable

production ML models

Required skills

SRE discipline, GPU workload management, Kubernetes orchestration, SLO definition and monitoring, incident response, capacity planning, observability (TTFT, tokens/sec), load testing, architecture judgment

Preferred skills

vLLM, TensorRT-LLM, Triton, Ray Serve, multi-GPU serving, quantization, batching, inference cost optimization, staff/lead SRE experience, open-source contributions to AI infrastructure

Technologies

Kubernetes, NVIDIA GPUs, vLLM, TensorRT-LLM, Triton, Ray Serve

Responsibilities

Maintain uptime, latency percentiles, and cost per million tokens for LLM serving; operate modern AI infrastructure with capacity planning and autoscaling; manage SLOs, observability, and incident response for AI workloads; harden the serving stack against traffic spikes and GPU operational issues; partner with product and engineering to ship AI features under real load

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗