CareerPlanSign in

Member of Technical Staff - Mid-Training Infra

San Francisco, CA💼 Full-time🗓 2026-03-24 → 2026-09-25

Core

Design, build, and operate large-scale GPU infrastructure for high-throughput model inference, mid-training workloads, and reinforcement learning pipelines.

Role type

Senior IC infrastructure engineer (GPU systems & distributed training)

Builds

High-performance inference platforms, synthetic data generation systems, and distributed RL training workflows

Domain

AI Infrastructure / Large Language Models / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Large-scale GPU system deployment, GPU performance optimization, distributed systems design, kernel-level optimization, model parallelism strategies, debugging distributed compute systems, inference framework expertise, RL pipeline infrastructure

Preferred skills

Experience with SGLang or Megatron, synthetic data pipeline experience, low-level runtime improvements

Technologies

SGLang, Megatron, GPU runtimes, distributed compute frameworks

Responsibilities

Design and operate GPU infrastructure for inference and mid-training, develop systems for synthetic data and RL pipelines, optimize throughput and latency for LLM workloads, build infrastructure for RL policy improvement loops, collaborate with research teams on distributed RL workloads, diagnose performance bottlenecks across GPU, networking, and distributed layers

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.