Senior Engineering Manager, ML Platform
Core
Lead the ML Platform Team to build foundational infrastructure for GPU compute, data storage, and training pipelines serving ML researchers and engineers.
Role type
Senior Engineering Manager, ML Platform (Player-coach)
Builds
Scalable GPU compute infrastructure, data storage/retrieval systems, and shared ML tooling frameworks
Domain
Machine Learning Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
GPU/distributed compute infrastructure, large-scale data storage and retrieval, data pipeline frameworks, architectural decision-making, team scaling, hands-on coding/debugging
Preferred skills
Compute orchestration frameworks (Kubernetes, Slurm, Ray), ML training workflows, dataset generation pipelines, feature stores, hiring ramp experience
Technologies
Kubernetes, Slurm, Ray, GPU clusters, TPUs
Responsibilities
Own strategy and roadmap for GPU compute infrastructure; oversee design of data storage and retrieval systems; lead development of shared libraries and frameworks; make key architectural decisions for compute orchestration and storage; actively participate in hiring and mentor engineers; define team processes for on-call and incident response
Seniority
Senior, hands-on IC with management responsibilities