AI Infrastructure Engineer, Model Serving Platform
Core
Design and build scalable, reliable platforms for serving Large Language Models (LLMs) to support internal research and external production systems.
Role type
Senior IC AI Infrastructure Engineer (Model Serving)
Builds
High-performance LLM serving platforms and internal capability discovery tools
Domain
AI Infrastructure / LLM Serving / Cloud Systems
Deliverable
infrastructure
Required skills
Large-scale backend system design, LLM serving fundamentals (rate limiting, token streaming, load balancing), container orchestration (Docker, Kubernetes), cloud infrastructure (AWS, GCP), Infrastructure as Code (Terraform)
Preferred skills
Modern LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, text-generation-inference)
Technologies
Python, Go, Rust, C++, Docker, Kubernetes, Terraform, AWS, GCP
Responsibilities
Build fault-tolerant, high-performance systems for LLM workloads; develop internal platforms for LLM capability discovery; collaborate with researchers to optimize models for production; conduct architecture reviews; develop monitoring and observability solutions; lead projects end-to-end
Seniority
Senior, hands-on IC