Member of Technical Staff, TPU Performance Engineering
Core
Build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure to make vLLM a first-class inference engine on Google TPUs.
Role type
Senior IC TPU performance engineer (inference systems)
Builds
Production-relevant model serving on TPU hardware with clear correctness, latency, and throughput benchmarks
Domain
AI inference systems, TPU hardware, compilers, and ML kernels
Deliverable
production ML models
Required skills
TPU workload optimization, JAX, XLA, Pallas, ML kernel optimization, performance profiling, benchmarking, TPU execution and memory behavior understanding
Preferred skills
vLLM, SGLang, TensorRT-LLM, XLA-based serving, compiler technologies (MLIR, LLVM), quantization methods (INT8, FP8, mixed precision)
Technologies
JAX, XLA, Pallas, vLLM, SGLang, TensorRT-LLM, MLIR, LLVM
Responsibilities
Build and optimize TPU backends and compiler integrations; optimize ML kernels and inference paths (attention, GEMM, sampling, KV cache); develop benchmarking infrastructure; profile and measure performance to guide optimization work
Seniority
Senior, hands-on IC