CareerPlanGet AI match score →

Research Engineer, Infrastructure, Inference

San Francisco💼 Full-time💰 $350,000–$350,000🗓 2026-05-04 → 2026-07-31

Core

Design, optimize, and scale infrastructure systems to enable high-performance, cost-effective, and reliable inference for large AI models.

Role type

Senior IC infrastructure research engineer (AI inference systems)

Builds

Scalable inference serving systems, orchestration frameworks, and compute fleets for AI models

Domain

Artificial Intelligence / Large Language Model Infrastructure

Deliverable

production ML models

Required skills

Deep learning frameworks (PyTorch, JAX), inference serving systems (SGLang, vLLM), distributed compute systems, GPU parallelism, hardware-aware optimizations, orchestration frameworks (Kubernetes, Ray, SLURM), codebase optimization, observability standards

Preferred skills

Experience with large-scale language models (hundreds of billions of parameters), open-source ML/systems contributions, improving research productivity via infrastructure design

Technologies

PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA

Responsibilities

Design and implement new techniques/tools/architectures to improve performance, latency, throughput, and efficiency; Optimize codebase and compute fleet (GPUs) to utilize hardware FLOPs, bandwidth, and memory; Extend orchestration frameworks for distributed inference, evaluation, and large-batch serving; Establish standards for reliability, observability, and reproducibility across the inference stack

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗