Member of Technical Staff, Training Infra Engineer
Core
Design and write high-performant, scalable software for training large models; bridge research and production by improving training infrastructure and cycles.
Role type
Senior IC training infrastructure engineer (distributed ML systems)
Builds
Production-ready training pipelines and tools for large-scale model development
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Python, JAX, PyTorch, XLA/MLIR, Kubernetes, Slurm, Ray, distributed training strategies, large-scale model training
Preferred skills
Publications at top-tier venues (NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)
Technologies
Kubernetes, Slurm, Ray, JAX, PyTorch, XLA, MLIR
Responsibilities
Design and write high-performant and scalable software for training; Improve training setup from infrastructure and codebase performance standpoint; Craft and implement tools to speed up training cycles; Research, implement, and experiment with ideas on supercompute and data infrastructure; Learn from and work with best researchers in the field
Seniority
Senior, hands-on IC