Senior ML Systems Engineer, Frameworks & Tooling
Core
Design and maintain the core components of a training framework for frontier-scale language models, enabling fast, reliable, and scalable training on large GPU clusters.
Role type
Senior ML Systems Engineer (Frameworks & Tooling)
Builds
Training framework, distributed training abstractions, monitoring/logging tooling, and reproducible large-scale run infrastructure.
Domain
Large-scale LLM training, distributed systems, HPC infrastructure
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed training, HPC systems, JAX internals, multi-node cluster orchestration, CUDA/NCCL debugging, containerized environments, performance engineering
Preferred skills
LLM training experience, ML framework contributions (PyTorch, JAX, DeepSpeed), evaluation/serving frameworks, data pipeline optimization
Technologies
JAX, Slurm, Ray, Kubernetes, CUDA, NCCL, Docker, Singularity/Apptainer, GB200/300, AMD, H200/100
Responsibilities
Build and own the training framework for large-scale LLM training; Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO); Improve training throughput and stability on multi-node clusters; Develop tooling for monitoring, logging, and debugging; Collaborate with infra teams on cluster and hardware configurations; Investigate and resolve performance bottlenecks across the ML systems stack
Seniority
Senior, hands-on IC