Performance Engineer
Core
Optimizing throughput and robustness of large-scale distributed machine learning systems.
Role type
Senior IC performance engineer (ML systems)
Builds
Low-latency high-throughput sampling systems, GPU kernels for low-precision inference, fault-tolerant distributed systems, custom load-balancing algorithms.
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Large-scale systems engineering, supercomputing scale experience, GPU/Accelerator programming, ML framework internals, OS internals, distributed systems design, kernel-level debugging, quantitative modeling of system performance
Preferred skills
Language modeling with transformers, high-performance computing
Technologies
GPUs, accelerators, transformers, containerized environments
Responsibilities
Identify novel systems problems in ML algorithms, develop systems to optimize throughput and robustness, implement low-latency sampling for large language models, write custom load-balancing algorithms, debug kernel-level network latency spikes, design fault-tolerant distributed systems
Seniority
Senior, hands-on IC