Senior and/or Principal Software Engineer- LLM Serving Performance
Core
Develop and optimize GPU kernels and runtime components to improve latency, throughput, and hardware efficiency for state-of-the-art LLM serving systems.
Role type
Senior/Principal IC software engineer (LLM serving performance)
Builds
Production LLM serving systems with optimized GPU kernels and runtime components
Domain
Artificial Intelligence / Large Language Model Inference / GPU Computing
Deliverable
production ML models
Required skills
GPU kernel optimization, C++ programming, Python programming, LLM inference parallelism, low-level architecture expertise, cross-team collaboration, technical ownership, AI-assisted development tool usage
Preferred skills
Expertise in vLLM or SGLang, experience with GPU compilation pipelines, track record of creating reusable platforms, ability to mentor technical leaders
Technologies
C++, Python, vLLM, SGLang
Responsibilities
Profile workloads to identify performance bottlenecks, integrate robust optimizations into production systems, collaborate with model/compiler/hardware teams, contribute to engineering standards and technical reviews
Seniority
Senior/Principal, hands-on IC with mentorship responsibilities
