Software Engineer - Training Infrastructure
Core
Architect and lead development of a global training platform for ML workloads, enabling research engineers to deploy, scale, and monitor high-performance training systems.
Role type
Senior IC infrastructure engineer (ML training stack)
Builds
Scalable scheduling, storage, networking, and observability systems for distributed ML training
Domain
AI/ML infrastructure, distributed systems, cloud-native platforms
Deliverable
production ML models
Required skills
Go, Kubernetes, distributed systems, observability, ML/AI workloads, MLOps
Preferred skills
distributed storage, Python, cloud providers (AWS/GCP/neo-cloud), workload orchestration (Temporal/Airflow), open-source training frameworks (PyTorch/Megatron/DeepSpeed)
Technologies
Go, Kubernetes, Temporal, Airflow, PyTorch, Megatron, DeepSpeed, NCCL, FSDP
Responsibilities
Design and architect scalable infrastructure systems for ML training; design a global training scheduler; design reinforcement learning systems and continuous learning pipelines; partner with developers to translate training requirements into technical solutions; drive reliability improvements and technical strategy; mentor junior engineers on infrastructure best practices
Seniority
Senior, hands-on IC with architectural leadership