Research Intern - 2027 Start
Core
Design and build large-scale container-based cluster management and orchestration systems for secure, cost-efficient ML platforms.
Role type
Research Intern (Systems & Infrastructure)
Builds
Cloud-native GPU and AI accelerator infrastructure, inference solutions, and production-ready systems.
Domain
Cloud infrastructure, distributed systems, AI/ML platforms, hardware acceleration.
Deliverable
production ML models | infrastructure
Required skills
large-model inference, distributed and parallel systems, high-performance networking, resource management, scheduling, request routing, monitoring, orchestration, Docker, Kubernetes, Go, Rust, Python, C++
Preferred skills
Kubernetes or Ray cluster management, GPU orchestration, workload scheduling, scaling, production isolation, CUDA, vLLM, SGLang, TensorRT-LLM, AWS, Azure, GCP, SageMaker, Azure ML, Vertex AI, Ray, DeepSpeed, PyTorch, distributed training or inference platforms
Technologies
vLLM, SGLang, TensorRT-LLM, Kubernetes, Ray, Docker, CUDA, PyTorch, DeepSpeed
Responsibilities
Design and build large-scale container-based cluster management and orchestration systems; Architect cloud-native GPU and AI accelerator infrastructure; Develop inference solutions using LLM engines; Integrate systems research into production systems; Write maintainable, testable, scalable, production-ready code; Collaborate across teams on systems research, distributed infrastructure, and hardware acceleration.
Seniority
Intern