Machine Learning Performance Engineer
Core
Design and implement techniques to optimize large-scale GPU and CPU workloads for research teams, ensuring efficient use of cutting-edge compute infrastructure.
Role type
Senior IC machine learning performance engineer
Builds
Optimized compute stack, reference implementations, libraries, and tools for large-scale ML computation
Domain
High-performance computing, distributed systems, machine learning infrastructure
Deliverable
production ML models
Required skills
Profiling and benchmarking distributed workloads, Python, C++, CUDA, deep learning frameworks (PyTorch), data structures and algorithms, parallel programming on heterogeneous systems, Linux OS fundamentals (scheduling, memory management, NUMA, networking, filesystems), HPC schedulers, Kubernetes-based workload orchestration, performance monitoring tools (nsys, ncu, eBPF)
Preferred skills
Experience with large-scale training and inference workloads
Technologies
PyTorch, CUDA, Kubernetes, Linux, nsys, ncu, eBPF
Responsibilities
Collaborate with researchers and engineers to understand compute challenges and design optimized solutions; Profile, benchmark, and tune large-scale training and inference workloads; Develop reference implementations, libraries, and tools to improve job efficiency and reliability; Collaborate with systems and platform teams to evolve the compute stack; Influence long-term platform and infrastructure decisions
Seniority
Senior, hands-on IC