Student Researcher - 2026 Start
Core
Design and optimize large-scale distributed training systems, reinforcement learning frameworks, and foundation-model inference performance for heterogeneous hardware.
Role type
PhD-level research engineer (ML systems & distributed computing)
Builds
Distributed training systems, RL training frameworks, inference optimization tooling, compiler/runtime optimizations for GPUs/accelerators
Domain
Machine Learning Systems, Distributed Computing, High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Python, C++, distributed computing, machine learning systems, performance optimization, GPU programming, compiler technologies
Preferred skills
large-scale ML systems, open-source ML systems contributions, performance tooling, publications in ML systems or distributed systems
Responsibilities
Design and optimize large-scale distributed training systems; Contribute to reinforcement learning training frameworks; Improve foundation-model inference performance; Develop compiler or runtime optimizations; Perform system-level performance analysis and profiling; Build tooling and automation for developer productivity
Seniority
PhD candidate / Researcher