Senior ML Systems Engineer, Inference
Core
Lead end-to-end LLM inference performance optimization to ensure the fastest and most cost-efficient serving for a million+ developers.
Role type
Senior IC ML Systems Engineer (Inference)
Builds
High-performance LLM serving runtimes, configurations, and measurement tooling for production GPU deployments.
Domain
AI Infrastructure / LLM Inference Systems
Deliverable
production ML models
Required skills
vLLM, SGLang, Python, LLM inference performance tuning, quantization, speculative decoding, distributed serving, GPU profiling, benchmarking
Preferred skills
CUDA/Triton kernel tuning, open-source contributions, multi-node GPU systems, high-speed networking
Technologies
vLLM, SGLang, Python, CUDA, Triton
Responsibilities
Define and build tooling for rigorous inference performance measurement (throughput, latency, cost); Profile and diagnose bottlenecks across the serving stack from scheduling to kernels; Optimize serving efficiency for large models on single and multi-node GPU deployments; Implement production-ready runtimes and defaults; Collaborate with product and infrastructure teams; Evaluate and adopt innovations from the inference ecosystem.
Seniority
Senior, hands-on IC