Lead Software Engineer, Model Serving Platform
Core
Architect and lead the development of a high-performance, next-generation model serving platform for multimodal AI foundation models.
Role type
Senior IC Lead Software Engineer (Model Serving Platform)
Builds
High-performance execution runtimes, distributed inference systems, and Python APIs for real-time AI applications.
Domain
AI Infrastructure / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
C++, Python, CUDA/HIP, Kubernetes/Ray, distributed systems design, LLM inference mechanics, performance profiling, system-level debugging
Preferred skills
ML systems engineering, distributed GPU scheduling, open-source inference engines (vLLM, Sglang, TRT-LLM), ROCm, large-scale MLOps infrastructure
Technologies
C++, Python, CUDA, HIP, Kubernetes, Ray, vLLM, Sglang, TRT-LLM, ROCm
Responsibilities
Lead technical direction and architecture decisions for the model serving platform; build core serving components including execution runtimes and distributed inference systems; develop high-performance GPU kernels and memory-optimized runtimes; collaborate with ML researchers to productionize multimodal models; mentor engineers through code reviews and design discussions; drive performance profiling and observability across the inference stack.
Seniority
Senior, hands-on IC with leadership responsibilities