Training: ML Framework Engineer
Core
Designing, implementing, and optimizing the core distributed machine-learning training runtime to accelerate research experiments and enable frontier-scale model runs.
Role type
Senior IC ML Framework Engineer (distributed systems)
Builds
High-performance, fault-tolerant training frameworks and distributed process management for large-scale GPU runs.
Domain
AI research infrastructure / Distributed systems / High-performance computing
Deliverable
production ML models
Required skills
Python, distributed systems, performance optimization, profiling, software engineering, supercomputer architecture
Preferred skills
experience with small-scale ML experiments, system design, reducing complexity
Technologies
Python, GPU clusters
Responsibilities
Apply latest techniques to achieve hardware efficiency for training runs, profile and optimize the training framework, work with researchers to enable next-generation model development
Seniority
Senior, hands-on IC