Machine Learning Systems Engineer, RL Engineering
Core
Build and maintain the algorithms and infrastructure for training AI models (RLHF) to support researchers in creating reliable, steerable AI systems.
Role type
Senior ML Systems Engineer (Reinforcement Learning)
Builds
High-performance distributed training systems and tools for LLM finetuning
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
Software engineering, distributed systems, Python, LLM finetuning algorithms (RLHF), system profiling, instrumentation, debugging training pipelines
Preferred skills
High performance computing, large scale distributed systems, LLM training experience
Technologies
Python
Responsibilities
Profile reinforcement learning pipelines to identify improvements; Build systems to launch and monitor training jobs; Modify finetuning systems for new model architectures; Implement instrumentation to resolve performance bottlenecks (e.g., GIL contention); Diagnose and fix training slowdowns; Implement new training algorithms proposed by researchers
Seniority
Senior, hands-on IC