CareerPlanSign in

Software Engineer, Inference

Redwood City, CA💼 Full-time🗓 2026-07-24 → 2026-09-26

Core

Design and operate large-scale inference systems to serve Luma's generative AI models across thousands of machines, optimizing GPU utilization and meeting strict SLOs.

Role type

Senior IC systems engineer (large-scale model inference)

Builds

High-throughput inference engine, scheduling systems, deployment pipelines, and internal observability tooling for model workflows.

Domain

Generative AI / Large-scale ML Systems / Cloud Infrastructure

Deliverable

production ML models

Required skills

Python, system architecture, model serving, Kubernetes, Linux, Docker, queue management, traffic control, fleet management, CI/CD

Preferred skills

RDMA (RoCE, InfiniBand, NVLink), high-performance ML systems (100+ GPUs), CUDA, FFmpeg

Technologies

PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, Redis, S3-compatible storage

Responsibilities

Integrate new model architectures into the inference engine; build scheduling systems to optimize expensive GPU resources; automate and maintain inference services for maximum uptime; manage and scale deployments across clusters and hardware providers; build tooling to profile and track inference job lifetimes.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.