AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
Core
Drive performance optimization for large-scale foundation model training on TPUs, focusing on efficiency, throughput, and scalability.
Role type
Staff ML Infrastructure Engineer (Pre-training)
Builds
High-performance TPU kernels and distributed training systems for foundation models
Domain
AI/ML Infrastructure, High-Performance Computing (HPC)
Deliverable
production ML models
Required skills
Python, distributed systems, parallel computing, performance optimization, TPU architecture, JAX, XLA, collective communication, kernel development (Pallas/Triton/CUDA)
Preferred skills
Advanced degree, GPU accelerator experience, PyTorch, large-scale foundation model training optimization
Technologies
TPU, JAX, XLA, Pallas, Triton, CUDA, PyTorch, ICI/Fabric
Responsibilities
Profile and optimize JAX/XLA workloads across compute, memory, and communication; Develop and optimize high-performance TPU kernels for attention and MoE; Optimize distributed training techniques and sharding strategies; Research and implement new techniques across the JAX, XLA, and TPU stack; Develop performance profiling, benchmarking, and automated tuning capabilities; Lead complex technical projects and mentor engineers
Seniority
Staff, hands-on IC with mentorship