AI Systems Research and Development Engineer – LLM Inference Systems & Optimization
Core
Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.
Role type
Senior IC systems engineer (LLM inference optimization)
Builds
Next-generation high-performance and intelligent inference systems for agentic enterprise workloads
Domain
AI Systems / Large Language Model Inference / High-Performance Computing
Deliverable
production ML models
Required skills
LLM inference system design, distributed AI systems, GPU systems, high-performance computing, CUDA/Triton programming, performance profiling (Nsight), system optimization, parallel decoding strategies, KV-cache management, model-system co-design
Preferred skills
AI-native engineering approaches, automated profiling and configuration search, open-source contribution
Technologies
vLLM, SGLang, TensorRT-LLM, CUTLASS, cuBLAS, cuDNN, Nsight Systems, Nsight Compute
Responsibilities
Design distributed inference strategies across GPUs and nodes; Develop efficient approaches for multi-model serving and dynamic resource management; Analyze and optimize GPU kernels and operators; Profile and benchmark end-to-end workloads to identify bottlenecks; Collaborate with model researchers and infrastructure teams to deploy innovations; Open-source and publish innovations.
Seniority
Senior, hands-on IC