Member of Technical Staff - Inference
Core
Design and optimize large-scale model serving systems to deliver high-performance inference for Grok to millions of users.
Role type
Senior IC distributed systems engineer (LLM inference)
Builds
High-concurrency, low-latency inference platforms serving billions of users
Domain
AI infrastructure / Large Language Model serving
Deliverable
production ML models
Required skills
C/C++ or Rust, GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM), distributed systems architecture, low-level GPU kernel optimization, quantization, speculative decoding, CI/CD infrastructure, benchmarking, reliability engineering
Preferred skills
Research on scaling test-time compute, RL rollout, model-hardware co-design
Technologies
vLLM, SGLang, Triton, TensorRT-LLM, C/C++, Rust, GPU kernels
Responsibilities
Architect scalable distributed infrastructure for model serving; Optimize latency and throughput under real production workloads; Build reliable high-concurrency serving systems; Benchmark and accelerate inference engines; Develop custom tools for tracing and fixing full-stack issues; Create robust CI/CD infrastructure; Accelerate research on scaling test-time compute and model-hardware co-design
Seniority
Senior, hands-on IC