AI Platform Engineer
Core
Remove compute-scaling bottlenecks for production LLMs by making frontier-model inference fast, efficient, reliable, and observable.
Role type
Senior IC LLM Inference Engineer (Systems/MLOps)
Builds
Production-grade LLM inference services serving products dependent on frontier models
Domain
Ecommerce, Large Language Models, High-Performance Computing, GPU Systems
Deliverable
production ML models
Required skills
Python, Go or Rust, LLM inference optimization, GPU architecture, CUDA/kernel programming, vLLM, Triton, PyTorch, SGLang, TensorRT, performance tuning, capacity planning, incident response, root-cause analysis
Preferred skills
Quantization, paging, kernel/runtimes improvements, speculative decoding, continuous batching, KV cache management
Technologies
vLLM, Triton, PyTorch, SGLang, TensorRT, CUDA
Responsibilities
Own production inference from handoff to serving including release engineering and incident response; Tune inference performance to reduce latency and increase throughput; Optimize runtimes and servers across heterogeneous GPU fleets; Benchmark and measure latency, throughput, and cost; Improve monitoring, tracing, and alerting for reliability; Apply and ship new optimizations like quantization and kernel improvements; Partner cross-functionally to translate requirements into performance SLOs
Seniority
Senior, hands-on IC