Machine Learning Systems Engineer
Core
Build and optimize core systems enabling researchers to train frontier models at scale, focusing on speed, cost, reliability, and throughput.
Role type
Senior IC machine learning systems engineer
Builds
High-performance distributed training pipelines and GPU kernels for large-scale model training
Domain
Autonomous vehicles / High-performance computing / Distributed systems
Deliverable
production ML models
Required skills
Python, PyTorch, GPU kernel development (Triton/CUDA), distributed training optimization, performance profiling, data pipeline engineering
Preferred skills
None stated
Technologies
PyTorch, Triton, CUDA, Nsight, PyTorch Profiler
Responsibilities
Profile and optimize data loading, gradient computation, and communication bottlenecks; Implement kernel fusion, sharding, and tiling optimizations; Design and maintain high-performance GPU kernels; Optimize robust data loading pipelines for maximum throughput
Seniority
Senior, hands-on IC