Member of Technical Staff, Integration/RL Team (Research Engineer)
Core
Designing and scaling machine learning algorithms and infrastructure for LLM post-training, specifically focusing on large-scale distributed reinforcement learning (RL) methods.
Role type
Senior IC machine learning engineer (distributed RL & post-training)
Builds
Production code for LLM post-training, distributed RL infrastructure, and research tooling
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models | infrastructure
Required skills
Software engineering, Python, JAX, PyTorch, XLA/MLIR, large-scale distributed training strategies, memory/speed profiling
Preferred skills
Kubernetes, Ray, post-training phase experience, ML/LLM/RL academic research
Technologies
Python, JAX, PyTorch, XLA, MLIR, Kubernetes, Ray
Responsibilities
Design and write high-performing, scalable software for training models; Develop new tools to support and accelerate research and LLM training; Coordinate with engineering and scientific teams to create an integrated post-training ecosystem; Craft and implement techniques to improve performance and speed up training cycles (SFT, offline preference, RL); Research, implement, and experiment with ideas on cluster and data infrastructure
Seniority
Senior, hands-on IC