Member of Technical Staff - Mid-Training Infra
Core
Design, build, and operate large-scale GPU infrastructure for high-throughput model inference, mid-training workloads, and reinforcement learning pipelines.
Role type
Senior IC infrastructure engineer (GPU systems & distributed training)
Builds
High-performance inference platforms, synthetic data generation systems, and distributed RL training workflows
Domain
AI Infrastructure / Large Language Models / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Large-scale GPU system deployment, GPU performance optimization, distributed systems design, kernel-level optimization, model parallelism strategies, debugging distributed compute systems, inference framework expertise, RL pipeline infrastructure
Preferred skills
Experience with SGLang or Megatron, synthetic data pipeline experience, low-level runtime improvements
Technologies
SGLang, Megatron, GPU runtimes, distributed compute frameworks
Responsibilities
Design and operate GPU infrastructure for inference and mid-training, develop systems for synthetic data and RL pipelines, optimize throughput and latency for LLM workloads, build infrastructure for RL policy improvement loops, collaborate with research teams on distributed RL workloads, diagnose performance bottlenecks across GPU, networking, and distributed layers
Seniority
Senior, hands-on IC