Software Engineer (Model Inference)
Core
Building and optimizing the end-to-end model inference stack to serve in-house and open-source AI models to tens of millions of users at low latency and high throughput.
Role type
Senior IC software engineer (ML inference & GPU systems)
Builds
High-throughput inference servers and optimized GPU serving frameworks
Domain
AI/ML inference, GPU-accelerated systems, interactive entertainment
Deliverable
production ML models
Required skills
GPU inference optimization, batching, quantization, CUDA kernel development, serving frameworks (vLLM, TensorRT, Triton), building software at scale
Preferred skills
Experience with LoRA testing and productionization, custom CUDA implementation
Technologies
vLLM, TensorRT, Triton, CUDA
Responsibilities
Design and ship high-throughput inference servers, optimize GPU utilization via batching and quantization, develop custom CUDA kernels, test and productionize models for millions of users
Seniority
Senior, hands-on IC