Senior ML Systems Engineer, Frameworks & Tooling
Core
Design and maintain the core components of a training framework for large-scale LLM training, connecting research ideas to thousands of GPUs.
Role type
Senior ML Systems Engineer (Frameworks & Tooling)
Builds
Distributed training abstractions, monitoring/logging/debugging tooling, and robust systems for reproducible large-scale runs.
Domain
Large-scale AI model training, distributed systems, HPC infrastructure
Deliverable
production ML models
Required skills
Large-scale distributed training, HPC systems, JAX internals, multi-node cluster orchestration, CUDA/NCCL debugging, containerized environments, performance engineering
Preferred skills
LLM training experience, ML framework contributions (PyTorch, JAX, DeepSpeed, Megatron), evaluation/serving frameworks, data pipeline optimization
Technologies
JAX, Slurm, Ray, Kubernetes, Docker, Singularity/Apptainer, CUDA, NCCL, GB200, AMD, H200/100
Responsibilities
Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO), improve training throughput on multi-node clusters, develop tooling for monitoring and debugging, collaborate with infra teams on cluster/hardware configurations, resolve performance bottlenecks, build robust systems for reproducible runs
Seniority
Senior, hands-on IC