Member of Technical Staff, Training Infra Engineer
Core
Design and write high-performant, scalable software for training frontier AI models and bridge the gap between research and production.
Role type
Senior IC training infrastructure engineer
Builds
Production-ready training pipelines and tooling for large-scale model training
Domain
AI/ML infrastructure, distributed computing, supercompute
Deliverable
production ML models
Required skills
Python, JAX, PyTorch, XLA/MLIR, Kubernetes, Slurm, Ray, distributed training strategies, large-scale model training
Preferred skills
Publications at top-tier venues (NeurIPS, ICML, ICLR, etc.)
Technologies
Kubernetes, Slurm, Ray, JAX, PyTorch, XLA, MLIR
Responsibilities
Design and write high-performant and scalable software for training; Improve training setup from infrastructure and codebase performance standpoint; Craft and implement tools to speed up training cycles; Research, implement, and experiment with ideas on supercompute and data infrastructure
Seniority
Senior, hands-on IC