AI Infra Engineer Graduate (Recommendation & LLM) 2027 Start (PhD)
Core
Build and optimize infrastructure for large-scale model training and online inference for TikTok's For You recommendation system and Large Language Models (LLMs).
Role type
PhD-level AI Infrastructure Engineer (LLM & Recommendation Systems)
Builds
Distributed training and inference systems, high-performance GPU computing, scalable LLM infrastructure for billions of daily recommendations
Domain
Internet / AI Infrastructure / Large Language Models
Deliverable
production ML models
Required skills
C++, Python, data structures, algorithms, computer systems, PyTorch, TensorFlow, Transformer architectures, LLMs
Preferred skills
LLM training/inference experience, distributed training (DP, TP, PP, FSDP, ZeRO), GPU programming (CUDA, Triton), LLM serving (KV Cache, Continuous Batching, FlashAttention), open-source contributions
Technologies
PyTorch, TensorFlow, CUDA, Triton
Responsibilities
Build and optimize infrastructure for large-scale model training and online inference; Develop distributed systems supporting large recommendation models and LLMs; Improve training and inference performance through GPU optimization and efficient communication; Collaborate with researchers to develop and deploy LLM training and serving solutions; Analyze system bottlenecks and implement performance optimizations
Seniority
PhD, Research/Engineering IC