LLM Engineer (Optimization)
Core
Develop AI systems that maximize LLM inference performance across server, edge, and on-device environments to serve real-world applications.
Role type
Senior IC LLM Infrastructure Engineer (Inference Optimization)
Builds
High-performance, low-latency, low-cost AI inference engines and runtimes for diverse hardware.
Domain
AI/ML Systems, LLM Infrastructure, GPU Computing
Deliverable
production ML models
Required skills
LLM Inference Optimization, GPU Architecture, CUDA Programming, Parallel Computing, Model Compression, Compiler Optimization, Python, C/C++
Preferred skills
vLLM, TensorRT-LLM, SGLang, llama.cpp, MLX, ONNX Runtime, Speculative Decoding, Prefill-Decode Disaggregation, KV Cache Compression, Expert Parallelism, Edge AI Optimization, Kubernetes AI Serving, LLM Fine-tuning, Distributed Training, Open Source Contributions, Top-tier Conference Publications
Technologies
vLLM, TensorRT-LLM, SGLang, llama.cpp, ONNX Runtime, MLX, CUDA, Triton, TensorRT, TVM, MLIR, NVIDIA Nsight Systems, NVIDIA Nsight Compute, py-spy, PyTorch, ONNX, NVIDIA H100, AMD GPU, Orin, Thor, Qualcomm, Apple Silicon
Responsibilities
Optimize LLM inference latency, throughput, and memory efficiency across various model structures and environments; Develop and optimize GPU/Accelerator-based LLM Inference Engines and Runtimes; Research and apply model compression techniques (Quantization, Pruning, Distillation) and Compiler Optimization; Develop optimization strategies for Edge AI and On-device LLM execution; Analyze performance and conduct benchmarks using profiling tools to resolve bottlenecks; Design Serving Architecture balancing model quality, performance, and cost.
Seniority
Senior, hands-on IC