CareerPlanGet AI match score →

Research Engineer, Infrastructure, Training Systems

San Francisco💼 Full-time💰 $350,000–$350,000🗓 2026-05-04 → 2026-07-31

Core

Design and build core systems enabling scalable, efficient training of large AI models for research and deployment.

Role type

Senior IC infrastructure research engineer (distributed training systems)

Builds

Distributed training stacks, high-performance optimization frameworks, and reusable libraries for large-scale model training

Domain

AI infrastructure / Distributed systems / Deep learning

Deliverable

production ML models

Required skills

Distributed systems design, High-performance computing (HPC), Deep learning frameworks (PyTorch, JAX), System architecture, Code optimization, Debugging complex codebases

Preferred skills

Experience with distributed training for large models, Open-source ML infrastructure contributions, Research productivity improvement

Technologies

PyTorch, JAX, XLA, Megatron-LM, DeepSpeed

Responsibilities

Design and optimize distributed training systems across thousands of GPUs, Develop high-performance optimizations for throughput, Build reusable frameworks for training reproducibility and scalability, Establish system reliability and security standards, Collaborate with researchers on scalable infrastructure, Publish technical reports and open-source libraries

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗