Research Intern, Inference (Winter 2027)
Core
Design and implement cross-layer optimizations for distributed inference systems, focusing on compiler-aware strategies, KV cache design, and large-scale serving architectures for foundation models.
Role type
Research Intern, Inference Systems
Builds
Efficient, scalable serving systems for large foundation models
Domain
AI Infrastructure / High-Performance Systems / Machine Learning
Deliverable
production ML models
Required skills
Machine Learning fundamentals, Deep Learning frameworks (PyTorch, JAX), Python programming, Transformer architectures, Distributed systems concepts
Preferred skills
CUDA programming, Model optimization techniques, Hardware acceleration approaches, Open-source contributions, Publications at top ML/systems conferences
Technologies
PyTorch, JAX, CUDA
Responsibilities
Design and conduct rigorous experiments to validate hypotheses, Document findings in scientific publications and blog posts
Seniority
Intern