Senior AI Engineer, SRE & LLM Infrastructure
Core
Ensuring the reliability, scalability, and cost-efficiency of large language model (LLM) serving in production environments.
Role type
Senior SRE / Production Engineer specializing in AI Infrastructure
Builds
High-availability LLM serving stacks on GPU-accelerated infrastructure
Domain
Artificial Intelligence / Large Language Models / Cloud Infrastructure
Deliverable
production ML models
Required skills
SRE discipline, GPU workload management, Kubernetes orchestration, SLO definition and monitoring, incident response, capacity planning, observability (TTFT, tokens/sec), load testing, architecture judgment
Preferred skills
vLLM, TensorRT-LLM, Triton, Ray Serve, multi-GPU serving, quantization, batching, inference cost optimization, staff/lead SRE experience, open-source contributions to AI infrastructure
Technologies
Kubernetes, NVIDIA GPUs, vLLM, TensorRT-LLM, Triton, Ray Serve
Responsibilities
Maintain uptime, latency percentiles, and cost per million tokens for LLM serving; operate modern AI infrastructure with capacity planning and autoscaling; manage SLOs, observability, and incident response for AI workloads; harden the serving stack against traffic spikes and GPU operational issues; partner with product and engineering to ship AI features under real load
Seniority
Senior, hands-on IC