AI Infrastructure Engineer Intern (Algorithm Infrastructure) - 2027 Start (PhD)
Core
Building inference infrastructure for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems to enable distributed serving, heterogeneous scheduling, and low-latency inference at massive scale.
Role type
PhD intern, AI infrastructure engineer (algorithm infrastructure)
Builds
Next-generation inference systems for large-scale online traffic, including global scheduling across heterogeneous compute resources, high-concurrency load balancing, and efficient batch formation
Domain
Artificial Intelligence, Large-scale Model Serving, Distributed Systems
Deliverable
production ML models
Required skills
System design for high-concurrency environments, Performance optimization, Production system development, CUDA programming, Triton programming, Knowledge of large-model architectures (MoE, attention mechanisms, multimodal fusion)
Preferred skills
Asynchronous scheduling, Resource pooling, Load balancing in distributed microservice systems
Technologies
CUDA, Triton, TP, EP, DP
Responsibilities
Build and evolve next-generation inference systems for large-scale online traffic; Optimize distributed inference for 200B+ models and complex multimodal models through TP, EP, DP, and related strategies; Develop high-performance kernels for frontier model architectures; Explore AI-driven infrastructure for inference systems
Seniority
Intern, PhD level