CareerPlanGet AI match score →

AI Infrastructure Engineer, Model Serving Platform

New York, NY💼 Full-time💰 $180,000–$180,000🗓 2026-07-09 → 2026-07-31

Core

Design and build scalable, reliable platforms for serving Large Language Models (LLMs) to support internal research and external production systems.

Role type

Senior IC AI Infrastructure Engineer (Model Serving)

Builds

High-performance LLM serving platforms and internal capability discovery tools

Domain

AI Infrastructure / LLM Serving / Cloud Systems

Deliverable

infrastructure

Required skills

Large-scale backend system design, LLM serving fundamentals (rate limiting, token streaming, load balancing), container orchestration (Docker, Kubernetes), cloud infrastructure (AWS, GCP), Infrastructure as Code (Terraform)

Preferred skills

Modern LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, text-generation-inference)

Technologies

Python, Go, Rust, C++, Docker, Kubernetes, Terraform, AWS, GCP

Responsibilities

Build fault-tolerant, high-performance systems for LLM workloads; develop internal platforms for LLM capability discovery; collaborate with researchers to optimize models for production; conduct architecture reviews; develop monitoring and observability solutions; lead projects end-to-end

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗