Member of Technical Staff - GPU Performance Engineer
Core
Design and ship custom CUDA kernels to optimize AI model training and inference pipelines for low latency and minimal memory usage.
Role type
Senior IC GPU Performance Engineer
Builds
Custom CUDA kernels, PyTorch pipeline extensions, and performance benchmarks
Domain
AI/ML infrastructure, GPU computing, high-performance computing
Deliverable
production ML models
Required skills
Custom CUDA kernel development, GPU architecture knowledge (memory hierarchy, warps, tensor cores), low-level profiling (Nsight Systems/Compute), C/C++ programming
Preferred skills
CUTLASS experience, Triton kernel development, PyTorch custom op integration, benchmark harness construction
Technologies
CUDA, PyTorch, Nsight Systems, Nsight Compute, C++, C, CUTLASS, Triton
Responsibilities
Write high-performance GPU kernels for novel model architectures, integrate kernels into PyTorch pipelines, profile and optimize training/inference workflows, build correctness tests and numerics checks, maintain performance benchmarks and guardrails, collaborate with researchers to ship speedups
Seniority
Senior, hands-on IC
