Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)
Core
Build and maintain high-performance training and inference systems for frontier AI research, enabling large-scale distributed RL and recursive self-improving agents.
Role type
Senior IC infrastructure engineer (AI research systems)
Builds
Distributed job orchestration, data pipelines, evaluation frameworks, and hardened RL/post-training libraries for scientific domains.
Domain
Frontier AI research, recursive self-improvement, large-scale distributed reinforcement learning
Deliverable
production ML models
Required skills
Large-scale distributed systems, Python, PyTorch or JAX, LLM training/inference technologies (SGLang, vLLM, verl, Megatron, OpenRLHF), system architecture trade-offs
Preferred skills
Systems programming (Rust or C++), CUDA kernel development, building reliable research tools
Technologies
SGLang, vLLM, verl, Megatron, OpenRLHF, PyTorch, JAX, CUDA
Responsibilities
Contribute to end-to-end training and inference infrastructure; build systems for large-scale experiments; implement and harden RL libraries; close recursive loops for agent-driven training; collaborate with core research teams
Seniority
Senior, hands-on IC