Distributed Training & Inference Optimization Engineer (LLM) - GPU Optimization Department (GPUOD)
Core
Maximize performance, efficiency, and scalability of LLM training and inference workloads on GPU clusters.
Role type
Senior IC GPU optimization engineer (LLM)
Builds
High-performance, cost-efficient AI infrastructure for large-scale LLM training and inference
Domain
Cloud computing, GPU-accelerated machine learning, LLMs
Deliverable
production ML models
Required skills
GPU-accelerated ML training & inference optimization, distributed training frameworks, LLM inference optimization, performance profiling, CUDA programming, system-level acceleration
Preferred skills
FlashAttention, PagedAttention, LoRA, speculative decoding, Kubernetes for GPU workloads, open-source ML framework contributions
Technologies
PyTorch, DeepSpeed, FSDP, Megatron-LM, vLLM, TensorRT-LLM, Triton, SGLang, NCCL, CUDA, Nsight, TensorBoard, KubeFlow, Volcano
Responsibilities
Optimize LLM training frameworks to maximize GPU utilization and reduce training time; Profile and optimize distributed training bottlenecks; Implement and tune inference optimizations for low-latency, high-throughput LLM serving; Collaborate with infrastructure teams to improve GPU cluster scheduling and fault tolerance; Develop benchmarking tools to measure training throughput and inference latency; Research and apply cutting-edge techniques to optimize LLM performance