Member of Technical Staff | ML Systems
Core
Build and operate the infrastructure for distributed machine learning training, model serving, and governance, enabling researchers to ship reliable models at scale.
Role type
Senior IC ML Systems Engineer (Distributed Systems & GPU)
Builds
High-performance distributed training infrastructure, experiment tracking systems, model registries, and governed release pipelines.
Domain
Machine Learning Infrastructure / Distributed Systems / GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Distributed machine learning training, CUDA kernel development, GPU performance optimization, columnar data formats (Lance, Arrow), experiment tracking, model lineage and versioning, Python development, multi-node workload management, software engineering principles for production systems
Preferred skills
Graph neural networks, graph sampling, multi-cloud GPU infrastructure (SkyPilot), model governance in regulated environments
Technologies
CUDA, Ray, PyTorch Distributed, Lance, Arrow, CSR/CSC, SkyPilot
Responsibilities
Build high-performance CUDA kernels for graph neural networks; Evolve distributed sampling and training infrastructure; Develop systems for efficient training and serving of ML models at scale; Define data formats and own data materialization pipelines; Establish data contracts with dataset teams; Build experiment tracking and evaluation infrastructure; Own model registry, lineage, and versioning; Define and operate release gates for governed model releases; Ensure full auditability of model releases; Optimize training throughput and GPU utilization; Build internal platform capabilities with a product mindset
Seniority
Senior, hands-on IC