Research Intern, Inference (Summer 2027)
Core
Design and implement cross-layer optimizations for distributed inference, compiler-aware strategies, and novel inference-time computation to lower cost and latency of large foundation models.
Role type
Research Intern, Inference Systems
Builds
Efficient, scalable serving systems for large foundation models
Domain
AI Infrastructure / High-Performance Systems / Deep Learning
Deliverable
production ML models
Required skills
Machine Learning fundamentals, Deep Learning frameworks (PyTorch, JAX), Python programming, Transformer architectures, Distributed systems, Compiler-aware optimization
Preferred skills
CUDA programming, Hardware acceleration, Model optimization techniques, Open-source contributions, Publications at top ML/systems conferences
Technologies
PyTorch, JAX, CUDA, Transformer architectures
Responsibilities
Design and conduct rigorous experiments to validate hypotheses, Document findings in scientific publications and blog posts
Seniority
Intern