Software Engineer, Inference
Core
Design and operate large-scale inference systems to serve Luma's generative AI models across thousands of machines, optimizing GPU utilization and meeting strict SLOs.
Role type
Senior IC systems engineer (large-scale model inference)
Builds
High-throughput inference engine, scheduling systems, deployment pipelines, and internal observability tooling for model workflows.
Domain
Generative AI / Large-scale ML Systems / Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, system architecture, model serving, Kubernetes, Linux, Docker, queue management, traffic control, fleet management, CI/CD
Preferred skills
RDMA (RoCE, InfiniBand, NVLink), high-performance ML systems (100+ GPUs), CUDA, FFmpeg
Technologies
PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, Redis, S3-compatible storage
Responsibilities
Integrate new model architectures into the inference engine; build scheduling systems to optimize expensive GPU resources; automate and maintain inference services for maximum uptime; manage and scale deployments across clusters and hardware providers; build tooling to profile and track inference job lifetimes.
Seniority
Senior, hands-on IC