Staff Software Engineer - GenAI Performance and Kernel
Core
Design, implement, and optimize high-performance GPU kernels for GenAI inference stacks, managing trade-offs between hardware efficiency and generality.
Role type
Staff Software Engineer (GenAI Performance and Kernel)
Builds
Production ML inference systems powered by optimized GPU kernels
Domain
Artificial Intelligence / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
GPU kernel development (CUDA, Triton, OpenCL, LLVM IR, assembly), GPU/accelerator architecture knowledge, advanced optimization techniques (tiling, fusion, auto-tuning), numerical stability reasoning, profiling and debugging tools, distributed inference systems
Preferred skills
Published research in systems/ML performance venues, custom accelerator/FPGA experience, sparsity or model compression techniques
Technologies
CUDA, Triton, OpenCL, LLVM IR, cuBLAS, cuDNN, CUTLASS, oneDNN, Nsight, NVProf, perf, vtune
Responsibilities
Lead design and implementation of core compute kernels (attention, MLP, softmax, memory management); Drive performance roadmap via vectorization, tensorization, and scheduling; Integrate optimizations with higher-level ML systems; Build profiling and verification tooling; Lead root-cause analysis on inference bottlenecks; Establish coding patterns for cross-backend portability; Mentor engineers on performance engineering
Seniority
Staff, hands-on IC with mentorship
