CareerPlanGet AI match score →

Member of Technical Staff - Inference

Palo Alto, CA💼 Full-time💰 $180,000–$180,000🗓 2026-05-26 → 2026-07-31

Core

Design and optimize large-scale model serving systems to deliver high-performance inference for Grok to millions of users.

Role type

Senior IC distributed systems engineer (LLM inference)

Builds

High-concurrency, low-latency inference platforms serving billions of users

Domain

AI infrastructure / Large Language Model serving

Deliverable

production ML models

Required skills

C/C++ or Rust, GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM), distributed systems architecture, low-level GPU kernel optimization, quantization, speculative decoding, CI/CD infrastructure, benchmarking, reliability engineering

Preferred skills

Research on scaling test-time compute, RL rollout, model-hardware co-design

Technologies

vLLM, SGLang, Triton, TensorRT-LLM, C/C++, Rust, GPU kernels

Responsibilities

Architect scalable distributed infrastructure for model serving; Optimize latency and throughput under real production workloads; Build reliable high-concurrency serving systems; Benchmark and accelerate inference engines; Develop custom tools for tracing and fixing full-stack issues; Create robust CI/CD infrastructure; Accelerate research on scaling test-time compute and model-hardware co-design

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗