ML Infrastructure Engineer
Core
Build scalable infrastructure for LLM post-training, RL, evaluation, inference, and agentic development workflows to directly impact model learning and product quality.
Role type
Senior ML Infrastructure Engineer
Builds
Distributed RL/post-training pipelines, data control systems, inference stacks, and agentic development environments
Domain
AI Safety, Large Language Models, Reinforcement Learning, Distributed Systems
Deliverable
production ML models
Required skills
Distributed RL/post-training systems design, Deep learning frameworks (PyTorch/JAX), Python (concurrency/async), Distributed GPU workload debugging, System profiling tools, Inference stack expertise, System metrics to model behavior reasoning
Preferred skills
Open-source contributions to ML infra, High-bar AI infra/research experience, Custom training framework ownership, Agentic coding systems usage, GPU cluster orchestration, Systems-level languages (Rust/C++/CUDA/Go)
Technologies
PyTorch, JAX, vLLM, SGLang, TensorRT-LLM, Dynamo, Kubernetes, Slurm, Ray, NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, EFA
Responsibilities
Build robust RL and post-training pipelines; Design data control systems for training data flow; Tune training and inference end-to-end for high throughput; Investigate infrastructure impact on learning dynamics; Build infrastructure for model iteration and reproducibility; Work on inference infrastructure affecting post-training loops; Build and improve agentic development environments
Seniority
Senior, hands-on IC