CareerPlanSign in

Senior ML Systems Engineer, Inference

🌐 Remote💼 Full-time💰 $150,000–$150,000🗓 2026-09-25 → 2026-09-28

Core

Lead end-to-end LLM inference performance optimization to ensure the fastest and most cost-efficient serving for a million+ developers.

Role type

Senior IC ML Systems Engineer (Inference)

Builds

High-performance LLM serving runtimes, configurations, and measurement tooling for production GPU deployments.

Domain

AI Infrastructure / LLM Inference Systems

Deliverable

production ML models

Required skills

vLLM, SGLang, Python, LLM inference performance tuning, quantization, speculative decoding, distributed serving, GPU profiling, benchmarking

Preferred skills

CUDA/Triton kernel tuning, open-source contributions, multi-node GPU systems, high-speed networking

Technologies

vLLM, SGLang, Python, CUDA, Triton

Responsibilities

Define and build tooling for rigorous inference performance measurement (throughput, latency, cost); Profile and diagnose bottlenecks across the serving stack from scheduling to kernels; Optimize serving efficiency for large models on single and multi-node GPU deployments; Implement production-ready runtimes and defaults; Collaborate with product and infrastructure teams; Evaluate and adopt innovations from the inference ecosystem.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 813,000+ jobs from 20+ sources.