Senior Performance Engineer - Deep Learning
Core
Building and optimizing open-source Deep Learning libraries (PyTorch, JAX) and tools to accelerate AI model training and deployment on NVIDIA hardware.
Role type
Senior IC performance engineer (Deep Learning systems)
Builds
Optimized Deep Learning frameworks, Transformer Engine, and MLPerf benchmark submissions
Domain
AI/ML infrastructure, GPU computing, High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
C++, Python, parallel systems programming, GPU programming, code optimization, computer architecture knowledge
Preferred skills
PyTorch, JAX, LLM architectures, CUDA, OpenAI Triton, cuBLAS, cuDNN, cuSOLVER, multi-GPU/multi-node profiling
Technologies
PyTorch, JAX, CUDA, OpenAI Triton, CuTeDSL, Pallas, Transformer Engine, MLPerf
Responsibilities
Build and support Transformer Engine for LLM training; implement and optimize new Deep Learning models; collaborate on systems research for low-precision training and parallelism; contribute to community benchmarks; engage with open-source community and enterprise customers; influence hardware generation design.
Seniority
Senior, hands-on IC