Senior Software Engineer, AI Inference Systems
Core
Architect and implement high-performance AI inference stacks, optimizing GPU kernels and compilers to serve large-scale models with extreme efficiency.
Role type
Senior IC machine-learning systems engineer (inference)
Builds
High-performance inference frameworks (vLLM), GPU kernels, compiler infrastructure, and benchmarking tools
Domain
AI/ML Systems, High-Performance Computing, GPU Architecture
Deliverable
production ML models
Required skills
Python, C/C++, CUDA, distributed systems, parallel programming, GPU memory hierarchy, container orchestration (Kubernetes/Docker), profiling/debugging
Preferred skills
Go, Rust, ML compilers (Triton, MLIR/LLVM), speculative decoding, disaggregation techniques
Technologies
vLLM, SGLang, PyTorch, Nsight Systems, Nsight Compute, Docker, Kubernetes, Slurm, NCCL
Responsibilities
Profile and optimize inference framework (vLLM) with parallelism and disaggregation techniques; Develop and benchmark GPU kernels using fusion and autotuning; Build high-level DSLs and compiler infrastructure; Architect scheduling for containerized large-scale inference on GPU clusters; Contribute to MLPerf Inference benchmarking suite; Conduct and publish original research on ML Systems
Seniority
Senior, hands-on IC