Senior System Software Engineer - AI Performance and Efficiency Tools
Core
Develop internal profiling, analysis, debugging, and benchmarking tools for AI workloads running on large-scale GPU clusters to empower researchers and engineering teams.
Role type
Senior System Software Engineer (AI Performance and Efficiency Tools)
Builds
Profiling and analysis tools, debugging utilities, benchmarking and simulation technologies for AI systems and GPU clusters
Domain
AI infrastructure, GPU clusters, high-performance computing
Deliverable
production ML models | product features | infrastructure
Required skills
C++, Python, CUDA Programming, NCCL, distributed training and inference, GPU cluster job scheduling (Slurm/Kubernetes), storage and networking, system debugging, performance analysis
Preferred skills
Linux device drivers, compiler implementation, GPU/CPU architecture knowledge, continuous profiling platforms, large-scale AI job performance analysis
Technologies
PyTorch, TensorFlow, CUDA, NCCL, Slurm, Kubernetes, C++, Python
Responsibilities
Build internal profiling and analysis tools for AI workloads at large scale; Build debugging tools for common problems like memory or networking; Create benchmarking and simulation technologies for AI system or GPU cluster; Partner with HW architects to propose new features or improve existing features with real world use cases
Seniority
Senior, hands-on IC