CareerPlanSign in

AI Infrastructure Engineer Intern (Algorithm Infrastructure) - 2027 Start (PhD)

San Jose, United States of America💼 Internship🗓 2026-09-28

Core

Building inference infrastructure for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems to enable distributed serving, heterogeneous scheduling, and low-latency inference at massive scale.

Role type

PhD intern, AI infrastructure engineer (algorithm infrastructure)

Builds

Next-generation inference systems for large-scale online traffic, including global scheduling across heterogeneous compute resources, high-concurrency load balancing, and efficient batch formation

Domain

Artificial Intelligence, Large-scale Model Serving, Distributed Systems

Deliverable

production ML models

Required skills

System design for high-concurrency environments, Performance optimization, Production system development, CUDA programming, Triton programming, Knowledge of large-model architectures (MoE, attention mechanisms, multimodal fusion)

Preferred skills

Asynchronous scheduling, Resource pooling, Load balancing in distributed microservice systems

Technologies

CUDA, Triton, TP, EP, DP

Responsibilities

Build and evolve next-generation inference systems for large-scale online traffic; Optimize distributed inference for 200B+ models and complex multimodal models through TP, EP, DP, and related strategies; Develop high-performance kernels for frontier model architectures; Explore AI-driven infrastructure for inference systems

Seniority

Intern, PhD level

Sourced via tiktok · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.