CareerPlanSign in

Software Engineer - Training Infrastructure

San Francisco💼 Full-time🗓 2025-08-29 → 2026-09-25

Core

Architect and lead development of a global training platform for ML workloads, enabling research engineers to deploy, scale, and monitor high-performance training systems.

Role type

Senior IC infrastructure engineer (ML training stack)

Builds

Scalable scheduling, storage, networking, and observability systems for distributed ML training

Domain

AI/ML infrastructure, distributed systems, cloud-native platforms

Deliverable

production ML models

Required skills

Go, Kubernetes, distributed systems, observability, ML/AI workloads, MLOps

Preferred skills

distributed storage, Python, cloud providers (AWS/GCP/neo-cloud), workload orchestration (Temporal/Airflow), open-source training frameworks (PyTorch/Megatron/DeepSpeed)

Technologies

Go, Kubernetes, Temporal, Airflow, PyTorch, Megatron, DeepSpeed, NCCL, FSDP

Responsibilities

Design and architect scalable infrastructure systems for ML training; design a global training scheduler; design reinforcement learning systems and continuous learning pipelines; partner with developers to translate training requirements into technical solutions; drive reliability improvements and technical strategy; mentor junior engineers on infrastructure best practices

Seniority

Senior, hands-on IC with architectural leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.