ML Systems Engineer, ML Acceleration
Core
Building core systems to enable researchers to train frontier models at scale, focusing on speed, cost, reliability, and throughput.
Role type
ML Systems Engineer (ML Acceleration)
Builds
High-performance distributed training pipelines and GPU kernels for large-scale model training
Domain
Autonomous vehicles / Machine Learning Systems
Deliverable
production ML models
Required skills
Python, PyTorch, GPU kernel development (Triton/CUDA), performance profiling, distributed training optimization, data pipeline engineering
Preferred skills
N/A
Technologies
PyTorch, Triton, CUDA, Nsight, PyTorch Profiler
Responsibilities
Profile and optimize bottlenecks in data loading, gradient computation, and communication; optimize distributed training pipelines; design and maintain high-performance GPU kernels; optimize robust data loading pipelines
Seniority
Mid-level, hands-on IC