Research Engineer, Infrastructure, Training Systems
Core
Design and build core systems enabling scalable, efficient training of large AI models for research and deployment.
Role type
Senior IC infrastructure research engineer (distributed training systems)
Builds
Distributed training stacks, high-performance optimization frameworks, and reusable libraries for large-scale model training
Domain
AI infrastructure / Distributed systems / Deep learning
Deliverable
production ML models
Required skills
Distributed systems design, High-performance computing (HPC), Deep learning frameworks (PyTorch, JAX), System architecture, Code optimization, Debugging complex codebases
Preferred skills
Experience with distributed training for large models, Open-source ML infrastructure contributions, Research productivity improvement
Technologies
PyTorch, JAX, XLA, Megatron-LM, DeepSpeed
Responsibilities
Design and optimize distributed training systems across thousands of GPUs, Develop high-performance optimizations for throughput, Build reusable frameworks for training reproducibility and scalability, Establish system reliability and security standards, Collaborate with researchers on scalable infrastructure, Publish technical reports and open-source libraries
Seniority
Senior, hands-on IC