AI Infra效率工程专家 - TikTok研发
Core
Analyze LLM training/inference/RL/SFT workloads to identify bottlenecks and optimize AI runtime, model serving, and resource efficiency for TikTok's core business scenarios.
Role type
Senior AI Infrastructure Engineer (LLM Performance & Efficiency)
Builds
High-throughput, low-latency AI inference and training platforms serving billions of users
Domain
AI Infrastructure, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
Go, C++, Python, Java, Linux/OS internals, Distributed systems, GPU architecture, Model training/inference pipelines, Kubernetes, System design
Preferred skills
PyTorch, vLLM, SGLang, Nsight, Auto-tuning, Agent development, AI Coding
Technologies
Kubernetes, PyTorch, vLLM, SGLang, Nsight, Go, C++, Python, Java
Responsibilities
Analyze workload performance and resource usage to locate bottlenecks in compute, memory, communication, and storage; Optimize AI runtime and model serving for heterogeneous GPUs, model loading, and inference engines; Build efficiency governance platforms for capacity planning and resource optimization; Evaluate and implement cutting-edge AI Infra technologies in production.