CareerPlanSign in

Lead Software Engineer, Model Serving Platform

San Francisco💼 Full-time🗓 2025-12-06 → 2026-09-26

Core

Architect and lead the development of a high-performance, next-generation model serving platform for multimodal AI foundation models.

Role type

Senior IC Lead Software Engineer (Model Serving Platform)

Builds

High-performance execution runtimes, distributed inference systems, and Python APIs for real-time AI applications.

Domain

AI Infrastructure / High-Performance Computing / GPU Systems

Deliverable

production ML models

Required skills

C++, Python, CUDA/HIP, Kubernetes/Ray, distributed systems design, LLM inference mechanics, performance profiling, system-level debugging

Preferred skills

ML systems engineering, distributed GPU scheduling, open-source inference engines (vLLM, Sglang, TRT-LLM), ROCm, large-scale MLOps infrastructure

Technologies

C++, Python, CUDA, HIP, Kubernetes, Ray, vLLM, Sglang, TRT-LLM, ROCm

Responsibilities

Lead technical direction and architecture decisions for the model serving platform; build core serving components including execution runtimes and distributed inference systems; develop high-performance GPU kernels and memory-optimized runtimes; collaborate with ML researchers to productionize multimodal models; mentor engineers through code reviews and design discussions; drive performance profiling and observability across the inference stack.

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.