AI Infrastructure Engineer Graduate (Algorithm Infrastructure) - 2027 Start (PhD)
Core
Building and evolving next-generation inference systems for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems, focusing on distributed serving, heterogeneous scheduling, and low-latency inference.
Role type
PhD-level AI Infrastructure Engineer (Algorithm Infrastructure)
Builds
High-performance inference systems for 200B+ models and complex multimodal models
Domain
AI Infrastructure / Large-Scale Model Serving
Deliverable
production ML models
Required skills
System design for high-concurrency environments, Asynchronous scheduling, Resource pooling, Load balancing, Performance optimization, CUDA programming, Triton development, Heterogeneous compute scheduling, High-concurrency load balancing, Batch formation, Kernel efficiency, Distributed inference strategies (TP, EP, DP), MoE architecture support, Emerging attention mechanisms, Multimodal fusion layers, AI-driven infrastructure development, AI Agents for optimization, Deployment pipelines, Consistency validation, Intelligent operations
Preferred skills
null
Technologies
CUDA, Triton, TP, EP, DP, MoE, Heterogeneous compute
Responsibilities
Build and evolve next-generation inference systems for large-scale online traffic, Optimize distributed inference for 200B+ models through TP, EP, DP, and related strategies, Develop high-performance kernels for frontier model architectures, Explore AI-driven infrastructure for inference systems
Seniority
PhD, Research & Engineering