ML Infrastructure Engineer
Core
Build scalable infrastructure for LLM post-training, RL, evaluation, inference, and agentic development workflows to directly impact model learning and product quality.
Role type
Senior ML Infrastructure Engineer
Builds
RL/post-training pipelines, data control systems, training/inference stacks, agentic development environments
Domain
AI Safety, Large Language Models, Reinforcement Learning, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
distributed RL/post-training systems, deep learning frameworks (PyTorch/JAX), Python (concurrency/async), distributed GPU debugging, system profiling, inference stacks (vLLM/SGLang/TensorRT-LLM), system metrics analysis
Preferred skills
open-source contributions to ML infra, experience at high-bar AI infra teams (xAI/ByteDance/Prime Intellect), custom training framework ownership, agentic coding systems, GPU cluster orchestration (Kubernetes/Slurm/Ray), systems-level languages (Rust/C++/CUDA/Go)
Responsibilities
Design and maintain distributed RL/post-training pipelines; build data control systems for training data flow; tune training/inference end-to-end for high throughput; investigate infrastructure impact on learning dynamics; build model iteration infrastructure (experiments/evals/dashboards); develop agentic development environments; collaborate on tradeoffs and planning
Seniority
Senior, hands-on IC