Member of Technical Staff - Reinforcement Learning (Infrastructure), AGI Autonomy
Core
Design, build, and maintain systems for training and evaluating state-of-the-art agent models using large-scale reinforcement learning on LLMs.
Role type
Senior IC machine-learning infrastructure engineer (reinforcement learning)
Builds
Training infrastructure for large-scale RL on LLMs
Domain
Artificial Intelligence / Machine Learning Systems
Deliverable
production ML models
Required skills
Python, Java, C++, neural deep learning methods, large-scale optimization, system troubleshooting, distributed systems, GPU programming
Preferred skills
Megatron, vLLM, Ray, patent or publication experience at top-tier conferences
Technologies
Megatron, vLLM, Ray, GPUs
Responsibilities
Develop training infrastructure for efficient and robust large-scale RL; work across low-level ML systems, job orchestration, and data management; analyze and profile complex ML systems to address performance bottlenecks; conduct MLSys research to create new techniques and tooling
Seniority
Senior, hands-on IC