Staff Software Engineer - GenAI Performance and Kernel
Core
Design, implement, and optimize high-performance GPU kernels for GenAI inference stacks, managing trade-offs between hardware efficiency and generality.
Role type
Staff Software Engineer (GenAI Performance and Kernel)
Builds
Highly-tuned, low-level compute paths for ML inference
Domain
Generative AI / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
CUDA, Triton, OpenCL, LLVM IR, assembly, GPU architecture (warp structure, memory hierarchy, tensor cores), advanced optimization (tiling, blocking, vectorization, fusion, auto-tuning), numerical stability analysis, profiling tools (Nsight, NVProf)
Preferred skills
Published in systems/ML performance venues, custom accelerators/FPGA experience, sparsity/model compression techniques
Technologies
CUDA, Triton, OpenCL, LLVM IR, cuBLAS, cuDNN, CUTLASS, oneDNN, Nsight, NVProf
Responsibilities
Lead design and implementation of core compute kernels (attention, MLP, softmax, memory management); Drive performance roadmap via vectorization, tensorization, and fusion; Integrate kernel optimizations with higher-level ML systems; Build profiling and verification tooling; Lead root-cause analysis on inference bottlenecks; Establish coding patterns for cross-backend portability; Mentor engineers on lower-level performance
Seniority
Staff, hands-on IC with mentorship